Releases: SAY-5/failsafe
Release list
v5.0.2 — portable Grafana authentication
v5.0.2 — private observability and portable Grafana authentication
Prometheus and Grafana publish on localhost by default. Grafana's default local
dashboard is an anonymous Viewer: no dashboard writes, basic authentication, or
login form, and no known initial administrator password.
Set FAILSAFE_GRAFANA_REQUIRE_AUTH=1 to require authenticated Grafana on the same
standard localhost address, including on macOS. Supply a unique
FAILSAFE_GRAFANA_ADMIN_PASSWORD of at least 16 characters through the process
environment. Empty, short, or whitespace-only passwords stop authenticated
startup. The flag accepts only 0 or 1; invalid values stop startup too.
Non-loopback Grafana bindings always require authentication, even when the flag
is 0. The README now demonstrates authenticated local startup with a fresh
Compose project rather than an alternate loopback address.
A dedicated CI gate starts the actual Grafana 11.1.0 image through Compose and
tests provisioned resources, read-only anonymous access, authenticated local and
remote-policy modes, literal password handling, and unsafe-startup refusal. It
uses ephemeral loopback host ports and isolated disposable projects. It checks
the Prometheus datasource configuration, not live Prometheus query results.
Upgrade notes
- Existing anonymous localhost defaults are preserved; authenticated localhost
requires an explicit opt-in. - Stop an older stack if its fixed host ports conflict with a new one.
- Use a fresh Grafana database for the demonstration. Initial password settings
do not rotate existing accounts; audit and rotate those credentials explicitly. - Use TLS, firewall restrictions, and appropriate identity management before
network exposure. This demo is not a production security configuration. - Gateway traffic, resilience policies, and upstream networking are unchanged.
v5.0.1
The browser demo under web/ moves from Vite 5.4.21 to 7.3.6 and @vitejs/plugin-react from 4.7.0 to 5.2.0, which drops the esbuild 0.21.5 that Vite 5 carried, and web/package.json now declares node ^20.19.0 || >=22.12.0, the range both packages require. npm audit over web/package-lock.json reported two vulnerable packages at v5.0.0, esbuild (moderate, GHSA-67mh-4wv8-2f99) and Vite (high, three advisories including GHSA-fx2h-pf6j-xcff), and reports 0 vulnerabilities at this commit, where the demo type-checks, bundles with Vite 7.3.6 and passes its 43 self-check assertions. Those checks were run by hand until now: CI gained a web job that installs with npm ci under the Node 22 that web/.nvmrc names, type-checks the demo, runs the self check and bundles it on every push. The gateway itself is unchanged: 108 tests pass at this commit, and both chaos suites kill replicas under load with 0 client-visible failures.
v5.0.0: operator control plane
The gateway gains a bearer-token admin API under /admin, disabled until admin_token is configured (the shipped config reads FAILSAFE_ADMIN_TOKEN and treats an unset variable as disabled rather than a literal secret). GET /admin/upstreams lists every replica with health, breaker state, ejection, draining flag, concurrency limit and recent error rate; the eject (optional seconds), readmit, drain, undrain and reset-breaker actions act on one replica through the same pool paths the outlier detector uses. POST /admin/drain withdraws readiness while the gateway keeps serving, so an instance can leave a load balancer gracefully. New metrics for draining replicas and admin actions. 108 tests.
v4.0.0: canary routing and outlier ejection
Routes can name a canary subset of replicas and a traffic weight; that share of requests is served by the subset, the rest avoids it, a header forces either side, and canary traffic falls back to the stable replicas when no canary can take it. Every replica keeps a window of recent outcomes, and after each health-check round a replica whose error rate or mean latency stands out from its peers is ejected for an escalating cool-down, never beyond half the pool and never when the whole service is failing alike. New metrics for canary requests, ejections and the ejected gauge; the chaos summary reports ejections. 101 tests.
v3.0.0: request hedging and deadline propagation
Idempotent requests that run longer than the route's observed p95 (or a fixed hedge.after_ms) now get a second attempt on another replica; the first successful answer is relayed and the loser is cancelled without recording anything on its breaker. Requests can carry an end-to-end budget through X-Request-Timeout or X-Request-Deadline, and routes can set a default deadline_seconds: the budget clamps every attempt timeout, is forwarded to upstreams with the remaining time, and stops a retry whose backoff would end past the deadline, answering 504 instead. Four new metrics cover hedges fired and won, the live hedge delay and deadline failures, and the chaos summary prints them. 88 tests.
v2.0.0: adaptive concurrency limits
Every upstream replica now carries an adaptive in-flight limit that grows while calls complete within the replica's no-load latency and is cut back on timeouts, resets, 5xx or a window whose average latency spikes. The pool skips replicas that are at their limit, so a slow replica stops receiving new work while its siblings carry it, and a request is only shed with 503 and Retry-After when no healthy replica has a free slot. Three new metrics expose the limit, the in-flight count and the shed total, and the chaos summary prints shedding alongside retries and failovers. 75 tests.
v1.0.0
First stable release of the gateway. Token-bucket rate limiting with exact Retry-After, per-replica circuit breakers, retries with full-jitter backoff and failover across replicas, and replica discovery from the Kubernetes EndpointSlice API. The compose and kind chaos suites both pass with zero client-visible failures while upstream containers and pods are killed under load, and the suite ships with 62 tests.