v5.0.1
Two tests compared a regenerated artifact with the committed one exactly, and that does not hold across CPU architectures: regenerated from its seed on the x86-64 CI runner, tests/fixtures/replay_log.jsonl gets distance_km 8.5 on the row where the committed file records 8.499, and the distance_km standard deviation in artifacts/manifest.json comes out 4.990105 against the recorded 4.990104, while an arm64 Mac reproduces the fixture byte for byte and the statistics exactly. test_fixture_is_reproducible_from_the_seeded_dataset and test_manifest_carries_training_stats_for_every_input now compare through assert_reproduces in tests/conftest.py, which holds each number to a tolerance declared per field: 0.001 for distance_km and 0.0001 for traffic_index, one unit in the last place modelgate/model/data.py rounds them to, 0.000001 for every figure modelgate/model/stats.py rounds to six decimals, and 0.02 for the two eta fields derived from those inputs. Keys, field order, line count, integers, strings and booleans still have to match exactly, and a decimal with no tolerance declared for it or for a container above it is refused, so a new field cannot pass unchecked. tests/test_artifact_tolerance.py pins the comparison: a difference inside the tolerance passes in either direction, two units in the last place, a drifted integer, a flipped boolean, a missing key or a changed shape fail, and a number outside its tolerance is reported with the field, the difference and the tolerance.
The README, the assert_reproduces docstring and the test comments that explain this comparison now describe it as the artifacts and the code show it. They had the direction reversed, saying the committed fixture records 8.5 where the runner regenerates 8.499; they named torch's vectorised exp as the cause, which no artifact isolates; they called every tolerance one unit in the last place, which the two eta fields are not; and they said a number with no declared tolerance must match exactly, where the code refuses it outright. The test pinning that refusal is renamed test_a_decimal_without_a_declared_tolerance_is_refused_even_when_equal. tests/measure_eta_shift.py makes the basis of the 0.02 eta tolerance reproducible: moving distance_km and traffic_index by one unit in their last place together, in every sign combination over the 300 fixture records, shifts the reference eta by at most 0.009381 minutes and the prediction v1 serves by at most 0.002720. The model and its artifacts are unchanged, so artifacts/manifest.json still records the 5.0.0 build that trained them.
Two timing budgets are widened. test_healthy_canary_stays_active, which failed on the CI runner when its healthy v2 canary was rolled back, now raises latency_floor_ms to 250 ms: a latency rollback needs the candidate's p95 to be over twice the primary's and more than that floor above it, and test_slow_canary_rolls_back_on_latency still proves the rule by injecting a 4 ms delay into the candidate. The batching test's max_wait rises from 5 ms to 50 ms with its assertions unchanged.
The browser demo under web/ moves from Vite 5.4.21 to 7.3.6 and @vitejs/plugin-react from 4.7.0 to 5.2.0 and declares node ^20.19.0 || >=22.12.0; npm audit over web/package-lock.json reported esbuild (moderate, GHSA-67mh-4wv8-2f99) and Vite (high, three advisories including GHSA-fx2h-pf6j-xcff) at v5.0.0 and reports 0 vulnerabilities at this commit, where the demo type-checks, bundles and passes its 15 self-check assertions, and a new CI web job runs npm ci under the Node 22 that web/.nvmrc names, the type-check, the self check and the bundle on every push. 116 tests at this commit, passing on the x86-64 CI runner and on an arm64 Mac, and the load test with shadow, pre-warm, a 0.1 canary and a mid-run swap drops 0 of 1800 requests.