Releases: CleverCloud/sozu-gateway
Release list
v0.5.0
sozu-gateway v0.5.0
A hardening release. It adds no new route type. Instead, it closes the ways one tenant's object, one metrics scrape or one data-plane restart could leave a shared gateway serving the wrong routes, or none, while every probe stayed green. Several fixes change what an operator sees: metrics series, collision winners, refused inputs and memory under load. That is why this is a minor bump. 15 commits since v0.4.0.
Highlights
- Known tenant inputs that froze the shared instance are now isolated. Translation is all-or-nothing, so a tenant input Sōzu rejected used to fail every reconcile. Until the object was deleted, no route, endpoint or certificate change was applied for any tenant. Three kinds of input are now refused on the object that carries them, using Sōzu's own loading and parsing checks plus a key/certificate pairing check where it can be verified: a
tls.keythat does not load, or that detectably does not belong totls.crt(a mismatched key used to load and fail every handshake instead) (InvalidCertificate); a regex path Sōzu cannot compile, including a literalExactorPrefixpath long enough to exceed the regex engine's compiled-size limit (InvalidPathRegex); and a hostname the apiserver accepts but Sōzu's IDNA processing refuses (InvalidHostname). A fourth case is arbitrated on the key Sōzu stores a route under: two HTTP routes that look distinct to the gateway, but that Sōzu stores under one key because a regex path contains;, are settled by oldest claimant, and the loser is reported withRouteCollision. The hostname and;cases were reproduced live. Before the fix, each one froze the shared instance. Now the object is refused on its own and the other routes keep converging. - Contested routes go to the oldest claimant, Ingress included. Ingress collision arbitration used to depend on the backend cluster-id order across namespaces. A contested route key now goes to the object with the oldest
creationTimestamp, thennamespace/name, for Ingress and HTTPRoute alike. - A Sōzu restart is handled end to end. The persisted shadow now records the command-socket identity of the Sōzu it was applied to, and is resumed only against that same Sōzu. When the controller detects a restart,
/readyzdrops, so the Pod leaves the Service until the full re-apply succeeds. A 2 s liveness tick notices a restart even on an idle socket. It also retries a reconcile that failed on a transient socket error, instead of waiting for an unrelated watch event. The shadow now advances as soon as the socket apply succeeds, before the post-apply identity check and the status writes, so a failing apiserver can no longer hold it back. - The recorded state stays honest. The command protocol has no request ids. The connection is now dropped after every channel error, including one on the retry, so a late reply can no longer be read as the next request's acknowledgement. A failed teardown is tolerated only for recognised already-absent answers, with each worker's answer checked on its own. A batch the controller has given up on stops between two requests.
- Default
/metricsscrapes can no longer wedge Sōzu. A scrape used to ask each worker for all of its per-cluster metrics in one message. On Sōzu 2.2.1, a worker whose message exceedsmax_command_buffer_sizerequeues that message forever. The worker then serves no traffic and takes no commands, while every probe stays green. With the chart's buffers, a local Sōzu 2.2.1 still answered at 2,000 clusters that had seen traffic, and both workers wedged at 2,500. Treat that as an order of magnitude, not a limit. Scrapes now ask for proxy-wide series only; per-cluster series are opt-in. - Sōzu's buffer pool matches its connection limit. Each worker allowed 10,000 connections but kept Sōzu's default pool of 1,000 buffers. Every open HTTP/1 connection or TCP session holds two buffers, idle keep-alives included, so a worker stopped near 500 connections while the Pod stayed Ready. With two workers, about 1,000 idle keep-alive clients made every fresh request fail. At the new default,
sozu.maxBuffers: 20000, fresh requests were still served with 6,000 connections held, and with 1,200 held on the test cluster. - Sticky sessions actually pin clients.
sozu.io/sticky-sessions: "true"never pinned anyone, because Sōzu matches a returning cookie against a sticky id that no backend carried. Each backend of a sticky Service now gets an opaque id, identical on every replica and across restarts. Before the fix, 8 requests replaying the cookie reached 3 Pods; now 10 of 10 reach one. The cookie no longer reveals the Service's namespace, name or Pod IPs. - Path rules are ready for Sōzu's next router. Upstream
mainanchors every regex path rule at both ends (47eb07c, marked breaking). On a data-plane bump, that would have turned mostPrefixandExactroutes into 404s. Generated rules are now full-span, so they mean the same on 2.2.1 and on its successor; the full-span rewrite leaves the earlier regex rules' matches on 2.2.1 unchanged.pathType: Exactnow compiles to an anchored regex instead of Sōzu'sEquals.Equalsmissed any request that carried a query string, and it could not be removed once installed, so a deletedExactroute kept serving.
Behaviour changes (read before upgrading)
Every one of these, and the downgrade procedure, is in docs/UPGRADING.md.
- Roll the complete gateway Pods. A file in the old shadow format is ignored once.
Exactrules already loaded asEqualscannot be removed by any request. Unchanged routes keep their old rule spelling in Sōzu, where a later removal would miss them. Ahelm upgradeto 0.5.0 replaces the Pods anyway, because the controller image and the rendered Sōzu config both change. Restarting only the controller container is not enough. /metricsexports proxy-wide Sōzu series only. By default, every series labelledcluster_iddisappears:sozu_requests, routed status counters such assozu_http_status_5xx, backend response and connection times, and backend availability. Alerts built on them now read "no data", not zero.metrics.perCluster: true(--metrics-per-cluster) brings them back, with the wedge risk above. If you enable it on a gateway that routes more than a few hundred Services, split the Services across Gateway instances.- Sōzu's buffer memory can grow with concurrency. With the defaults, the pool can now reach about 630 MiB (
workerCount × maxBuffers × 16 KiB), where it used to stop at about 32 MiB. TLS state comes on top of that. Idle memory is unchanged. Ifresources.sozusets a memory limit, size it for your real peak or lowersozu.maxBuffers. Raisesozu.maxBuffersif many clients speak HTTP/2. The newsozu.maxConnectionsPerIpdefaults to0(unlimited). Its429rejections are counted per Service port, so only the access logs ormetrics.perCluster: trueshow them. - A liveness tick probes Sōzu every 2 s.
controller.sozuProbeSecs(--sozu-probe-secs) sends oneStatusround-trip per tick;0disables it. A worker that restarts on the same socket no longer resets the shadow, because the main process re-feeds its state to that worker. - Ingress collisions follow the oldest claimant. An install that relied on the old cluster-id order may see a different object win a contested
host+path. - Some inputs are now refused. A
tls.keythat is a valid key but does not belong totls.crtstops loading; every handshake for those names already failed. A regex match Sōzu cannot compile is dropped and reported (InvalidPathRegex): an HTTPRoute staysAccepted, and an Ingress path is skipped. A hostname that fails IDNA is dropped from its route. A Gateway listener getsAccepted: False/UnsupportedValuewhen its hostname, without a leading*., fails IDNA; the frontend names it emits are checked separately. None of these names ever served traffic. After upgrading, checkkubectl get eventsand route status. - User regex paths should be full-span now. They are handed to Sōzu verbatim, so their meaning will change with Sōzu's next release. For example, write
^/api(?:[/?](?-u:.*))?$; an Ingress path can anchor behind/{0}^. A literal;that must not clash with a method can be written[;]or\x3B. A user regex that overlaps aPrefixorExactpath on the same host may swap winners. A regex that spelled the previous internal form of such a rule is now a second, overlapping route; rewrite it as the plain path it meant. - Inferred certificate names are stored lowercase, which is how rustls presents the SNI. Before, no handshake selected a mixed-case SAN. The change takes effect when the Pods roll.
--watch-timeout-secsdefaults to60in the binary, not only in the chart. An install whose values predatecontroller.watchTimeoutSecs, or a--reuse-valuesupgrade from one, was silently unbounded and no longer is.0now opts out in the provisioner as well, which used to read it as60. The chart renders an explicit0. Values of295and above are refused at startup instead of freezing every cache behind a green/readyz.- The first reconcile re-sends every backend of each sticky Service once. Sōzu applies the update in place. A client whose connection stays open may keep its old cookie until the connection closes, and is then pinned again. Two sticky Services on different paths of one host share the cookie and can overwrite each other's pin.
- Teardown failures are no longer all tolerated. A removal Sōzu refuses for a real reason now fails the reconcile, which is then retried from the unchanged shadow. The command-socket client's acknowledgement deadline per request attempt, and its socket write timeout, both drop from 30 s to 20 s.
Fixes
- The Ingress+TLS e2e suite now asserts what it observes instead of only printing it, so a...
v0.4.0
sozu-gateway v0.4.0
A feature release: a Gateway can now get its own data plane and its own address without a second list of names in Helm, the chart exposes as many HTTP and HTTPS listeners as the exposure table declares, header filters mean what the Gateway API says they mean, and the data plane moves to Sōzu 2.2.1. The replica default drops from three to two — because the difference was measured rather than assumed. 43 commits since v0.3.1.
Highlights
- One instance per Gateway, provisioned automatically. With
gatewayProvisioning.enabled: true, every Gateway whose GatewayClass names this controller gets a dedicated controller + Sōzu Deployment, Service and ConfigMap in the release namespace, plus the disruption budget and metrics resources the release enables. A separate provisioner owns the generated resources; installation and Gateway UIDs guard every update, cleanup and same-name recreation, and a foreign object with a colliding name is never adopted. The original Deployment keeps serving Ingress and stays the only writer of Ingress and GatewayClass status. There is no list of Gateway names to maintain, and adding a Gateway needs no Helm change. Off by default, because enabling it moves existing Gateway addresses. - Multiple HTTP and HTTPS listeners. Every
exposureentry of protocolHTTPorHTTPSis rendered as a Sōzu listener, a container port and a Service port, in both the default instance and the provisioned ones. The first HTTP and HTTPS entries keep serving Ingress and the health probes. - Header filters mean what they say.
RequestHeaderModifier/ResponseHeaderModifier:setreplaces the client's or backend's value,addappends,removedeletes.setused to append, so a value the client sent was never discarded. Emptyset/addvalues are rejected withAccepted: False/UnsupportedValue, because in Sōzu an empty value is a deletion. - Endpoint changes no longer wait behind the debounce. Referenced EndpointSlice changes interrupt the ordinary watch debounce and coalesce in their own bounded channel; an empty relist withdraws endpoints that left the cache. Measured on 20 endpoint changes and 1,600 fresh requests per run: median time from patch to the third consecutive new-backend response fell from 732 ms to 221 ms, with every request answered
200throughout. - A restarted Sōzu is detected even when it reuses its worker PIDs. A restarted container can come back with every PID unchanged and no routes; the controller kept its shadow and served
404indefinitely with/readyzgreen. The command socket's identity is now part of the generation. Measured: the previous build stayed at404for the full 180 s window, the fixed build recovered in 58 s with unchanged PIDs and no controller restart. - Sōzu 2.2.1 is the chart's default data plane. The protobuf schema, channel framing and routing diff are unchanged from 2.2.0; upstream fixes frontend validation and command failure handling, and redacts sensitive debug output.
- Two replicas by default, not three. Forcing the loss of one gateway Pod — what a rolling node replacement does, one worker at a time — cost a single replica 4.8% and 8.0% of fresh connections over two trials, and cost two and three replicas nothing. The third replica earned its keep only when a second loss overlapped the first (3.2% for two replicas, clean for three). It is worth its capacity if you want that margin:
replicaCount: 3.
Behaviour changes (read before upgrading)
Every one of these, and the downgrade procedure, is in docs/UPGRADING.md.
- Header
setreplaces instead of appending. A route that relied on the client's value surviving asetnow loses it. - HTTPRoute collisions have a stable winner. Among frontends sharing one route key: oldest
creationTimestamp, then alphabeticalnamespace/name, then first rule. Backend names no longer influence the choice, so a previously colliding route may change backend on the first reconcile — including mixed Ingress/HTTPRoute collisions, where the Ingress can now win over a newer HTTPRoute. - UDP flows are keyed by client IP and source port. Clients behind one address no longer receive each other's replies. Roll the complete gateway Pods: a controller-only restart with an unchanged shadow does not replace an installed cluster's UDP settings, and the rollout drops existing flows.
- A Gateway with
spec.infrastructure.parametersRefis rejected withAccepted: False/InvalidParameters, and contributes no routes or certificates. Earlier versions ignored the reference and served the routes without the requested configuration. - A Gateway with no accepted listener is no longer
Accepted. All listeners rejected reportsAccepted: False/ListenersNotValid; a mix reportsAccepted: True/ListenersNotValid. - Shared certificates keep the names a hostname-less listener infers, alongside the explicit hostnames of other listeners. Roll the gateway Pods when adopting this: Sōzu 2.2.1 acknowledges a same-fingerprint replacement without updating its SNI names, so a controller restart alone cannot repair an already loaded certificate.
- The Sōzu image changes to
2.2.1, which rolls the gateway Pods. If your values pinimage.sozu.tag, or you upgrade with--reuse-values, set--set image.sozu.tag=2.2.1to adopt it. replicaCountdefaults to2. Nothing to do unless you set it yourself, in which case your value still wins. A release adopting the new default rolls one Pod out; the budget and the drain keep that gap-free.- Enabling
gatewayProvisioningmigrates addresses, and neither direction is atomic. Enabling moves owned Gateways to new Services; disabling moves their routes back to the shared instance. Allow for LoadBalancer provisioning, watchProgrammed, and update DNS or clients before relying on the new addresses. The earlier, unreleasedgatewayInstancesvalue is rejected outright. NetworkPolicies and monitors that select only the default Deployment's labels need selectors for the generated Pods too (sozu.io/installation-uid,sozu.io/gateway-uid).
Fixes
- Layer-4 routes that lose a listener's socket stay
Accepted: True, count toward the listener'sattachedRoutes, and name the winning route and port in their ownAcceptedmessage, so the explanation survives the Warning Event's expiry; Gateway API CRDs installed after startup are noticed on the resync tick; a TCPRoute probe that closed its write side before the response arrived no longer fails a working route; scoped Gateway workers no longer watch Ingress and IngressClass cluster-wide for a cache they never read, and the provisioner's debounce can be interrupted by shutdown and resync. - A reproducible Gateway API conformance runner lives under
tests/conformance/, with explicit cluster selection and reports that keep their failures.
Known limitations
- Gateway API conformance. The last full, unconditioned run recorded for the default branch is 20/37 core, 1/3 extended on the
GATEWAY-HTTPprofile (v1.6.1). Several official v1.6.2 tests were measured passing individually on the fixes merged here — header modifiers, listener protocol rejection, multiple Gateways, non-TCP listener rejection — but no full run of this release is recorded. A 53/57 figure exists in the history; it was measured on an experimental integration carrying workarounds since withdrawn, and does not describe this release. The profile cannot fully pass on Sōzu (no weighted backend split, no header/query matching, no HTTP 500). Seedocs/E2E-RESULTS.md§6 anddocs/features.md. - SNI names of an already loaded certificate are not updated on a same-fingerprint replacement by Sōzu 2.2.1. Adding or removing a hostname-less listener, or changing explicit hostnames, while keeping the certificate needs a gateway Pod rollout until the native update tracked in #81 is available. Rotation to a different fingerprint is unaffected.
- Only routes with a resolved backend reference compete for a layer-4 socket, so an older route whose Service disappears falls back to the next claimant. Native path precedence and route update limitations are tracked in #80.
- NodePort address publication is not implemented. Provisioned
ClusterIPServices publish their internal address; a pending LoadBalancer publishes nothing. /readyzstill latches green on a controller that has gone blind. v0.3.1 bounded how long that lasts (controller.watchTimeoutSecs: 60); a watch-freshness signal is still the natural follow-up.URLRewriteand redirect path rewriting remain reported rather than wired: a literal$in a rewrite value makes Sōzu reject the frontend, and a path rewrite drops the query string.
Install
helm upgrade --install sozu-gateway \
oci://ghcr.io/clevercloud/sozu-gateway \
--version 0.4.0 --namespace sozu-system --create-namespace --waitArtifacts
- Image:
ghcr.io/clevercloud/sozu-gateway-controller:v0.4.0 - Chart:
oci://ghcr.io/clevercloud/sozu-gateway(0.4.0) - Data plane:
clevercloud/sozu:2.2.1, driven bysozu-command-lib=2.2.1
v0.3.1
sozu-gateway v0.3.1
A patch release with one subject. v0.3.0 shipped a known limitation: a controller whose Kubernetes watches stop delivering keeps programming Sōzu from a frozen cache, and nothing bounded how long that lasted. It is now bounded by default — and the workaround v0.3.0 suggested for it is withdrawn, because it was measured on the next upgrade and does not work. 5 commits since v0.3.0.
Highlights
- Watch blindness is bounded by default.
controller.watchTimeoutSecsdefaults to60. A replaced control plane leaves the connections its watches ride on open and silent rather than closing them: nothing errors, no event arrives, the reflectors stop advancing, and every reconcile keeps succeeding against a cache that has stopped moving. Sōzu then serves routes for backends that left minutes ago. - Measured, across three managed-cluster upgrades. With the previous default, 19.7%, 17.1% and 14.7% of fresh HTTP requests through the gateway failed for the duration of node replacement — every one an HTTP
504from Sōzu against a withdrawn backend, while bare TCP, DNS and the application's own Service stayed at0.000%,reconcile_failures_totalstayed0and/readyzstayed green throughout. WithwatchTimeoutSecs: 60, the same upgrades measured0.000%, in every window. - A second bound, off by default.
controller.kubeReadTimeoutSecsbounds the connection rather than the watch. It is kept as a knob and left off: it tears the connection instead of resuming from the storedresourceVersion, so it pays a watch error and a backoff every time it fires on a quiet watch, and the arm running it carried two unexplained defects across the measurement runs. - Two wrong comments corrected. One claimed the shared kube client keeps a ~295s read timeout.
kube::Configsetsread_timeout: Nonein every constructor, and that figure belongs to the watcher's idle timeout instead — a bound on watch streams and nothing else. That belief is what let this failure mode through in the first place.
Behaviour changes (read before upgrading)
Both are in docs/UPGRADING.md, with the downgrade path.
controller.watchTimeoutSecsdefaults to60. Ahelm upgradewith unchanged values rolls the Pod and changes how the controller talks to the apiserver. The cost is one watch reconnect per watch per minute; the watcher resumes from the storedresourceVersionrather than re-listing, so a quiet cluster pays one request per watched kind per minute and nothing else. Set it to0to restore the previous behaviour.- The v0.3.0 workaround is withdrawn. v0.3.0 advised rolling the gateway once the control plane had been replaced. An arm whose three controllers were restarted on exactly that event — each confirmed restarted and Ready, one Pod at a time — still lost 16.53%, indistinguishable from the untreated control. The fresh watches attach to an apiserver that is itself about to be withdrawn. Restarting is only safe once the new apiserver is serving, which the provider's step event does not tell you.
Known limitations
/readyzstill latches green on a controller that has gone blind, and nosozu_gw_controller_*series moves while it is blind — including the last-successful-reconcile timestamp, which keeps advancing because the reconciles genuinely succeed. This release bounds how long the blindness lasts; it does not make it visible. A watch-freshness signal and a metric for it are the natural follow-ups.- The nominal bound is not the measured one. Cutting controllers off from the apiserver and timing each to its first logged watch error gives 336–364s at the default and 117–132s at 60, against nominals of 295 and 65. Idle expiry logs only at
DEBUG, so those figures include the failed reconnect that follows it, and why they exceed their nominal bound is not established — only that they do. Size any staleness alert on the measured figures. - The root cause is upstream. An apiserver leaves without closing the connections pinned to it; the in-cluster
svc/kubernetesendpoint dropped the old address inside a 1.6s window that brackets the last events the watches ever received, and nothing closed those sockets. This release only bounds how long a controller suffers it. - Gateway API conformance is unchanged from v0.3.0: the
GATEWAY-HTTPprofile scores 20/37 core, 1/3 extended on v1.6.1, and cannot fully pass on Sōzu. Seedocs/E2E-RESULTS.md§6 anddocs/features.md.
Install
helm upgrade --install sozu-gateway \
oci://ghcr.io/clevercloud/sozu-gateway \
--version 0.3.1 --namespace sozu-system --create-namespace --waitArtifacts
- Image:
ghcr.io/clevercloud/sozu-gateway-controller:v0.3.1 - Chart:
oci://ghcr.io/clevercloud/sozu-gateway(0.3.1) - Data plane:
clevercloud/sozu:2.2.0, driven bysozu-command-lib=2.2.1
v0.3.0
sozu-gateway v0.3.0
A feature release: raw TCP and UDP now route through the Gateway API, the Gateway API types move to v1.6.1, the data plane moves to Sōzu 2.2.0, and the chart's defaults change shape — the component whose loss is immediately user-visible no longer ships as a single replica. 45 commits since v0.2.0.
Highlights
- Layer 4 through the Gateway API.
TCPRouteandUDPRouteon aprotocol: TCP/UDPGateway listener replace the cluster-globaltcp/udp-servicesConfigMaps. The listeners are user-defined, so unlike HTTP/HTTPS they are carried in the IR and created + activated dynamically over the command socket. Two routes contesting a socket are settled in the builder by oldestcreationTimestampthennamespace/name, so one tenant's conflict can no longer fail the whole reconcile. - Gateway API v1.6.1. Redirect
hostname/path/porttargets are honoured;allowedRoutes.namespaces.from: Selectoris evaluated against a Namespace label index instead of failing closed; eachparentRefgets its ownstatus.parents[]entry keyed on the whole ref (sectionNameandportincluded), so a route naming one Gateway per listener no longer collapses into a single entry whoselastTransitionTimemoves on every pass.GatewayClass.status.supportedFeaturesis published. pathType: Prefixmatches on element boundaries, not a raw string prefix./foocovers/foo,/foo?q=1and/foo/bar, and no longer/foobar. Sōzu matches path rules with the query string attached and does not anchor regexes, so a non-root prefix compiles to an anchored regex; the root/stays a plain prefix.- The data plane defaults to three replicas, one per node, with a hard hostname spread and a PodDisruptionBudget. A single replica has no one to fail over to: while its Pod is rescheduled the Service holds no endpoint at all.
- Sōzu is drained on shutdown instead of waiting out the kill. Sōzu registers no signal handler, and a signal with no handler is discarded at PID 1 — SIGTERM did nothing, and every rollout and node drain paid the full grace period and bought nothing. A preStop hook now asks over the command socket; the grace period grows to 40s.
- Prometheus
/metricsis served by default, because nobody enables metrics after the outage that needed them.
Behaviour changes (read before upgrading)
Every one of these, and the downgrade procedure, is in docs/UPGRADING.md.
- The
tcp/udp-servicesConfigMaps are removed. Existing L4 entries must be migrated toTCPRoute/UDPRouteobjects. A layer-4 port can never be443, and no bind may be privileged; both are rejected by the chart rather than left to a raw apiserver rejection. replicaCountdefaults to3, spread hard across nodes. With fewer nodes than replicas the surplus Pods stayPendingand ahelm upgrade --waitdoes not complete. Lower the count, or relax the placement.- The chart declares
kubeVersion: ">=1.27.0-0". The new defaults rely onmatchLabelKeys, which is a hard rejection under strict field validation before 1.27. - The backend connect budget drops from Sōzu's 3 seconds to 2. A backend that legitimately takes longer to accept starts being answered
504where it previously waited. The point is attribution, not recovery: a connect that times out gets no retry and no failover, so it is worth having it counted and named rather than consuming the caller's whole budget as an unexplained hang.sozu.timeouts.connectraises it. - Prometheus metrics are on by default. The scrape rides the same command socket as routing applies, and the endpoint is unauthenticated on a shared cluster.
- The grace period grows to 40s. What the hook drains is HTTP/1 exchanges in progress and HTTP/2 streams; layer-4 sessions, WebSockets, TLS handshakes in progress and idle keep-alives are cut once the delay elapses.
Fixes
- A hung API call no longer parks the reconcile loop for minutes; a duplicate L4 port is resolved in the builder instead of being left to fail translation; an L4 address held by another cluster is reported instead of retried forever; Gateway API CRDs installed after startup are noticed on the resync tick; a Gateway API enum member this build does not know no longer fails the parse; a route sharing no hostname with its listener is refused; a
RequestRedirectSōzu cannot honour is skipped rather than half-applied, and a redirect status it cannot emit is no longer served as a302; Warning Events reference the owning object by uid sokubectl describeshows them; an older shadow file stays readable; the L4 requests carry the SNI routing fields Sōzu 2.2.0 adds.
Known limitations
- A controller whose Kubernetes watches stop delivering keeps serving the routes it last saw.
watcherretries internally, so a broken watch surfaces as neither an item nor an error and the reflector simply stops advancing; the periodic resync does not repair it, because it rebuilds from the same caches. Measured across a managed-Kubernetes cluster upgrade: controllers that predated the control-plane replacement kept routing to backends removed minutes earlier whilereconcile_failures_totalstayed0and/readyzstayed green, costing about a fifth of new connections for the duration. Controllers started afterwards were unaffected, so a restart cures it — roll the gateway once the control plane has been replaced. - Gateway API conformance is a documented partial: the
GATEWAY-HTTPprofile scores 20/37 core, 1/3 extended on v1.6.1. The profile cannot fully pass on Sōzu (no weighted backend split, no header/query matching, no HTTP 500), so the recorded failures are not regressions. Seedocs/E2E-RESULTS.md§6 anddocs/features.md. URLRewriteand redirect path rewriting remain reported rather than wired: a literal$in a rewrite value makes Sōzu reject the frontend, and a path rewrite drops the query string.
Install
helm upgrade --install sozu-gateway \
oci://ghcr.io/clevercloud/sozu-gateway \
--version 0.3.0 --namespace sozu-system --create-namespace --waitArtifacts
- Image:
ghcr.io/clevercloud/sozu-gateway-controller:v0.3.0 - Chart:
oci://ghcr.io/clevercloud/sozu-gateway(0.3.0) - Data plane:
clevercloud/sozu:2.2.0, driven bysozu-command-lib=2.2.1
v0.2.0
sozu-gateway v0.2.0
A hardening release: a full-repo review (code + design) drove 46 commits closing four critical-class defects, a security batch, user-facing diagnostics, and reconcile-loop performance work. Routing behaviour is unchanged for well-formed configurations; several silent failure modes now fail closed and say why.
Highlights
- Reconciliation can no longer wedge or silently blackhole. Duplicate frontend adds after a missed ack are repaired (evict + reinstall) instead of aborting every reconcile forever; a restarted Sōzu container under a live controller is detected via its worker-PID generation and triggers a full re-apply (previously: indefinite 404s with a Ready pod); the shadow file is written atomically and no longer trusted when torn.
- One tenant can no longer break everyone. A corrupt TLS Secret is fully validated (X509 DER parse) in the builder and reported as a per-Secret problem instead of freezing cluster-wide routing convergence; host+path collisions across namespaces are detected and reported on the losing object instead of silently re-routing traffic.
- Users can finally see why. Problems publish as Warning Events on the owning Ingress/Gateway/HTTPRoute and appear in status condition messages with the actual detail (which Secret, which Service port, which listener). Operators get controller self-metrics (
sozu_gw_controller_*), including the last-successful-reconcile timestamp — the one-line staleness alert.
Behaviour changes (read before upgrading)
- TLS Secrets must be
type: kubernetes.io/tls. The controller now watches only that type (memory + blast-radius bound). Opaque Secrets carryingtls.crt/tls.keyare no longer seen.kubectl create secret tlsand cert-manager output are unaffected. - Gateway listeners must declare the advertised ports (default 80/443, configurable via
--gateway-http(s)-port; the chart wires them from the Service values). A mismatchedlistener.portis rejected withPortUnavailableinstead of silently landing on the static listeners. - Fail-closed instead of fail-open/silent:
allowedRoutes.namespaces.from: Selectoradmits nothing and is reported (previously admitted every namespace); a singlebackendRefwithweight: 0is rejected (previously received 100% of traffic);rule.timeoutsand per-backendRef filters are reported as unsupported;spec.defaultBackendis reported instead of silently ignored; hostless Ingress TLS entries and FQDN EndpointSlices are reported instead of half-handled. - Chart: default resource requests are now set for both containers (pods leave BestEffort QoS — scheduling may change); the ServiceAccount token is no longer mounted into the Sōzu container; the shared runtime volume is memory-backed and Sōzu's root filesystem is sealed (with a dedicated tmpfs /tmp for its worker-fork state handoff); L4 ConfigMap RBAC is namespaced; Gateway status writes are gated by
rbac.allowGatewayStatusWrites(defaulttrue); a PodDisruptionBudget renders whenreplicaCount > 1.
Security
- The pod's ServiceAccount token (cluster-wide Secret read) is projected into the controller container only — an exploit in the internet-facing proxy no longer holds API credentials.
- TLS private keys are stripped from the persisted shadow; certificate material never lands on node disk (tmpfs volume).
- CI/release installs
justfrom a pinned, checksum-verified release (no morecurl | sudo bash); e2e suites deploy their throwaway ttl.sh image by digest.
Fixes
- Backend property changes no longer delete the backend through command reordering; duplicate
(listener, fingerprint)certificates merge by fingerprint with SNI-name union (no moreReplaceCertificatechurn or lost SNI coverage); wildcard route × specific listener programs the hostname intersection (no over-matching); the cache-sync gate covers every reflector (no route flap on controller restart); socket writes/connects and/metricsscrapes are time-bounded;SOZU_GW_RESYNC_SECS=0disables resync instead of panicking; transient Gateway API discovery errors fail fast instead of silently locking Ingress-only mode; conflicting L4 entries for one port now fail explicitly instead of being programmed ambiguously.
Performance
- The builder borrows the reflector caches (
Arc) instead of deep-cloning every cached object each reconcile; EndpointSlice churn for services no route references no longer triggers rebuilds.
Known limitations
- The recorded Gateway API conformance score (16/33 core) predates the
Selectorfail-closed change; two recorded passes were artifacts of the fixed fail-open behaviour. A re-run is planned (see docs/E2E-RESULTS.md §6).
v0.1.1
Maintenance release.
Packaging: the controller image is now published as ghcr.io/clevercloud/sozu-gateway-controller — previously it shared the sozu-gateway repository with the Helm chart, collapsing both into one ghcr package. The chart keeps the sozu-gateway name and the helm install command is unchanged. No functional changes to the controller.
Install
helm upgrade --install sozu-gateway \
oci://ghcr.io/clevercloud/sozu-gateway \
--version 0.1.1 --namespace sozu-system --create-namespace --waitArtifacts
- Image:
ghcr.io/clevercloud/sozu-gateway-controller:v0.1.1 - Chart:
oci://ghcr.io/clevercloud/sozu-gateway(0.1.1)
v0.1.0
First release of sozu-gateway — a Kubernetes Ingress controller and API gateway built on the Sōzu reverse proxy. It watches Kubernetes objects, compiles them into a neutral intermediate representation, and hot-applies the minimal set of mutations to a co-located Sōzu over its command socket — no proxy restarts.
Highlights
- Ingress + TLS — exact + wildcard hosts,
Prefix/Exact/regex paths, SNI termination from Secrets with zero-gap rotation, pod-IP backends from EndpointSlices. - Gateway API (v1.2.1 standard channel) —
GatewayClass/Gateway/HTTPRoute/ReferenceGrant, withAccepted/Programmed/ResolvedRefsconditions and per-listener status. - HTTPRoute filters — request/response header edits and request redirects (scheme + status).
- Raw TCP/UDP (L4) forwarding via the chart's
l4.tcpServices/udpServicesmaps. - Opt-in Prometheus
/metrics— pulled from Sōzu over the command socket. - Idempotent hot reload — a single global reconcile applies only the delta; no proxy restart.
Install
helm upgrade --install sozu-gateway \
oci://ghcr.io/clevercloud/sozu-gateway \
--version 0.1.0 --namespace sozu-system --create-namespace --waitStatus & limits (pre-1.0)
Usable but pre-1.0 — APIs and defaults may change. Gateway API conformance is partial (official GATEWAY-HTTP profile, v1.2.1: core 16/33). The remaining gaps are Sōzu data-plane limits (no HTTP 500 on an invalid backendRef, no weighted backend split, no header/query-value matching, header set appends rather than replaces) and the single-LoadBalancer deployment model (catch-all collisions); URLRewrite is reported unsupported. ARM64 node pools are unsupported (the upstream clevercloud/sozu image is amd64-only). See docs/E2E-RESULTS.md and docs/features.md.
Artifacts
- Image:
ghcr.io/clevercloud/sozu-gateway:v0.1.0 - Chart:
oci://ghcr.io/clevercloud/sozu-gateway(0.1.0)