Skip to content

v5.1.0

Choose a tag to compare

@SAY-5 SAY-5 released this 27 Sep 00:49
· 30 commits to main since this release
de182ad

The Kubernetes rolling update proof now proves what it claims, and the load generator can survive its own run.

The generator had been killed at a 512Mi ceiling: it runs one virtual thread per driver and per ride, and the JDK HTTP client holds per connection buffers outside the heap, so the job now asks for 1Gi and says why. Its sends are bounded, at 1200 driver pings and 100 ride submissions in flight, and a send past the bound is skipped and counted rather than queued, so a target that falls behind cannot grow the set of live requests until the heap gives out. A ride is retried once on a transport failure, the way a position ping already was, because a keep alive connection closed by a draining pod is not an answer; a non 2xx status is distinguished from a transport failure and is not retried, and the first error causes are logged instead of swallowed. The script waits for the job to reach Complete or Failed rather than Complete alone, captures the generator's output while the pod is still alive, and reports a termination reason, so a dead generator fails the run in seconds instead of burning the timeout. The three Deployments are replaced one at a time, since nine service pods on one node starve each other, and decisions are asserted from the durable trip rows in the city shards rather than from the in process counters a replacement resets.

The verdict is a measured coverage gate rather than a green rollout status. The generator records the boundaries of its active load and samples successful sends once a second, the script records one UTC interval per Deployment, and scripts/verify-rollout-evidence.py requires every rollout interval to sit inside sustained load with a two second margin at either end, no sampled progress gap over five seconds, and at least ninety percent of the configured rate for both rides and pings across each interval. An HTTP error or a skipped ride submission fails the verdict; skipped driver pings fail it above one percent of the pings delivered and are reported either way, because that bound is the generator declining to queue a position write while a pod drains rather than a caller seeing the update. scripts/test_rollout_evidence.py drives the gate with fifteen cases, twelve of which must fail, including the summary of the run whose ninety six second rollout outlasted its sixty second load.

The README's evidence block is the output of the first hosted run to pass that gate, run 36282821355: three replacements of 31.7, 37.9 and 44.5 seconds, adding up to the 114 second rolling update the run reports, each with its own recorded interval and, over the slightly wider span the surrounding samples cover, a measured rate of 10.0 rides a second and 599 to 613 pings a second, all three inside a 180 second load window, 1804 rides submitted with no errors and none skipped, 115,200 position pings with no errors and none skipped, and 1804 of 1804 trip rows decided and matched. The paragraph above it explains the one number a reader is most likely to misread: the latency reservoir belongs to the matching pod that answered the last read, which has usually just taken over, so its sample leans on the rides that were waiting in the retry store when it did.