Skip to content

Releases: SAY-5/dispatchgrid

v5.2.0

Choose a tag to compare

@SAY-5 SAY-5 released this 29 Sep 03:53
5195398

A minor release that changes how the compose demo's figures are produced and quoted. The services, the wire format in common/, the migrations, the service manifests and the load generator's command line are untouched; make demo keeps its name and its stack, the generator's summary JSON keeps all 31 fields 5.1.0 wrote and adds eight, and each new environment variable has a default.

The load generator's summary now opens with a provenance object: the commit and machine passed in as RUN_COMMIT and RUN_MACHINE (unspecified when absent), the container's platform and JDK, the load window in UTC to the second, and its kernel's load average at either end of the load. It stores every figure its text prints, renders the text from those fields alone, stores it as summaryText, sets complete last, and prints the text and a SUMMARY_JSON line carrying the same map. The text gains commit, window, machine and load average lines, and the matching counters line now shows retried and dropped. LoadGenTest asserts, against a stub server, that the stored fields render the stored text exactly, that complete comes last and that the printed line equals the file.

make demo now runs scripts/compose-demo.sh, which starts the same stack, then runs the generator with --no-deps so that run --build no longer rebuilds and recreates the three services just before the load. It saves the SUMMARY_JSON line as demo-out/loadgen-summary.json, since the generator's own file is removed with its container, and writes demo-out/demo-run.json with the commit (suffixed -dirty on a modified tree), the host and Docker VM, the start and finish times, the host load average before and after, the stack's memory halfway through the load and the generator's exit code. A failed attempt keeps the last completed pair. The script needs git and a checkout, which the old compose commands did not. scripts/k8s-e2e.sh passes a commit and machine into the loadgen Job and refuses values outside a small character set.

scripts/patch-demo-summary.py writes the README block from those two files or refuses: an incomplete run or a nonzero generator exit, files that disagree on the commit or machine, a measured window outside the run's span, a modified tree, a commit other than HEAD, a missing or placeholder field, or a text line its fields do not render to. The README says what it cannot catch: a field edited together with its line, and the caption's load averages and memory, copied unchecked. With --check DIR it writes nothing and fails unless the README's block is, to the character, what the run committed in DIR renders to, under the same rules except that the run's commit need only be an ancestor of HEAD; the build job in CI runs it. scripts/test_demo_summary.py drives the patcher with 21 cases. scripts/demo-load-gate.sh, run before make demo, waits up to 45 minutes for a host load average below 8 and a Docker VM load average below 3 on two consecutive readings 30 seconds apart.

The README demo block captured at ae5dba8 on 2026-09-03 is gone; the 5.1.0 README already said the printed summary carried fields it lacked, and it named no clock window or load average. The block now comes from one make demo run at 1bc404c, 18:12:37 to 18:14:17 UTC on 2026-09-28, whose summary, run record and gate readings are committed in docs/demo-runs/1bc404c/; the readings end with two consecutive passes under those limits, and none of the code the run exercised has changed since. On Darwin arm64 with 10 CPUs and a 6 CPU, 7.7 GiB Docker VM, with host load 5.62 before and 4.68 after and container kernel load 2.18 at the start of the load and 5.75 at the end, it submitted 603 rides with no errors, retries or skips, decided and matched 603 of 603 trip rows, and measured p50 15 ms, p95 56 ms and p99 221 ms, with 1496 MiB in use across eight containers halfway through. The 5.1.0 claim that the stack fits a 2 GiB Docker VM is gone too.

Tests at the release commit, in CI run 36472860778: lint, 36 script cases, the demo block check, 63 Surefire unit tests and 14 Testcontainers integration tests, all passing, and the kind rolling update under load, whose rollout gate printed COVERAGE: PASS and RESULT: PASS.

v5.1.0

Choose a tag to compare

@SAY-5 SAY-5 released this 27 Sep 00:49
de182ad

The Kubernetes rolling update proof now proves what it claims, and the load generator can survive its own run.

The generator had been killed at a 512Mi ceiling: it runs one virtual thread per driver and per ride, and the JDK HTTP client holds per connection buffers outside the heap, so the job now asks for 1Gi and says why. Its sends are bounded, at 1200 driver pings and 100 ride submissions in flight, and a send past the bound is skipped and counted rather than queued, so a target that falls behind cannot grow the set of live requests until the heap gives out. A ride is retried once on a transport failure, the way a position ping already was, because a keep alive connection closed by a draining pod is not an answer; a non 2xx status is distinguished from a transport failure and is not retried, and the first error causes are logged instead of swallowed. The script waits for the job to reach Complete or Failed rather than Complete alone, captures the generator's output while the pod is still alive, and reports a termination reason, so a dead generator fails the run in seconds instead of burning the timeout. The three Deployments are replaced one at a time, since nine service pods on one node starve each other, and decisions are asserted from the durable trip rows in the city shards rather than from the in process counters a replacement resets.

The verdict is a measured coverage gate rather than a green rollout status. The generator records the boundaries of its active load and samples successful sends once a second, the script records one UTC interval per Deployment, and scripts/verify-rollout-evidence.py requires every rollout interval to sit inside sustained load with a two second margin at either end, no sampled progress gap over five seconds, and at least ninety percent of the configured rate for both rides and pings across each interval. An HTTP error or a skipped ride submission fails the verdict; skipped driver pings fail it above one percent of the pings delivered and are reported either way, because that bound is the generator declining to queue a position write while a pod drains rather than a caller seeing the update. scripts/test_rollout_evidence.py drives the gate with fifteen cases, twelve of which must fail, including the summary of the run whose ninety six second rollout outlasted its sixty second load.

The README's evidence block is the output of the first hosted run to pass that gate, run 36282821355: three replacements of 31.7, 37.9 and 44.5 seconds, adding up to the 114 second rolling update the run reports, each with its own recorded interval and, over the slightly wider span the surrounding samples cover, a measured rate of 10.0 rides a second and 599 to 613 pings a second, all three inside a 180 second load window, 1804 rides submitted with no errors and none skipped, 115,200 position pings with no errors and none skipped, and 1804 of 1804 trip rows decided and matched. The paragraph above it explains the one number a reader is most likely to misread: the latency reservoir belongs to the matching pod that answered the last read, which has usually just taken over, so its sample leans on the rides that were waiting in the retry store when it did.

v5.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 08 Sep 22:01

Trip timeline and pickup ETA. GET /rides/{id}/timeline returns every ride_events row for a trip from the shard that owns it, oldest first (requested, each retry with its radius and reason, the match with distance, latency and surge, cancellation or completion). The match now persists driver_distance_m and GET /rides/{id} turns it into pickupEtaSeconds while a driver is on the way, using rider.pickup-speed-mps (8 m/s by default). Full mvn verify green: 55 unit tests and 14 Testcontainers integration tests, topology IT at 2229 matches per minute.

v4.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 05 Sep 22:25

Rides that find no free driver no longer fail on the first pass. The request path is now a Processor API node with a changelogged pending-retries store and a wall-clock punctuator: attempt n waits n times the base backoff (0 s, 5 s, 15 s by default), each pass appends ride.retry to the trip timeline, and only the last failure moves the row to UNMATCHED and publishes RideUnmatched with the attempt count. The pending set follows the task to the standby replica through the changelog. Covered by TopologyTestDriver tests with a mocked wall clock and a Testcontainers scenario where a driver appears after the first pass. 51 unit tests and 14 integration tests.

v3.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 05 Sep 22:21

Trip lifecycle. POST /rides/{id}/cancel moves a REQUESTED or MATCHED trip to CANCELLED in the rider service and POST /drivers/{id}/trips/{ride}/complete publishes a completion; both travel on the new ride-lifecycle topic to the matching service, which applies the conditional row update (only the matched driver may complete) and releases the Redis claim so the driver is back in the pool immediately instead of at claim expiry. completed_at and cancelled_at columns, ride_events for every transition, completed/cancelled/lifecycleIgnored stats. 49 unit tests and 14 integration tests.

v2.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 05 Sep 08:21

Per-city, per-cell surge pricing computed in the matching service from ride requests and driver positions: trailing-window demand, TTL-decayed supply, a clamped multiplier stamped on every match and persisted on the trip row, GET /pricing/{city} and /quote, and surge gauges per city. Two matching replicas with a standby and static membership so a rolling update no longer rebalances the Streams group. 43 unit tests and 13 integration tests.

v1.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:56

First tagged release of the marketplace matching platform: rider request, driver location, and Kafka Streams matching services with atomic Redis claims, page-growth-then-widen candidate search, city-keyed MySQL shards, and zero-downtime rolling updates on Kubernetes. 32 unit tests and 13 Testcontainers integration tests cover routing, claims, the full topology, and the rollout script.