Skip to content

Releases: SAY-5/modelgate

v5.0.2 — private CLI startup

Choose a tag to compare

@SAY-5 SAY-5 released this 29 Sep 09:18
0c6408c

ModelGate v5.0.2

The installed modelgate serve command now defaults to 127.0.0.1 and checks
the same publication policy as the Make/Compose launcher before starting
Uvicorn. Non-loopback hosts refuse unset, empty, whitespace-only and known
development admin tokens. Explicit host, port and log-level options remain
available; configured tokens remain literal and an empty local token disables
the admin API.

This closes the CLI entry point missed by the previous startup hardening.
Package metadata, /healthz and OpenAPI now consistently identify 5.0.2.

Verification: 183 tests, Ruff lint/format and the locked dependency check passed
locally. A local 150 rps, 12-second run completed 1,800/1,800 requests with zero
drops during prewarming, shadow/canary traffic and a mid-run model swap.
Hosted CI passed
both test/load and web jobs for PR head 428fda076a2aed87cd2bd4e684245ca497ea8214.

This is still a development demo. Admin tokens protect /admin/*, not
predictions, metrics or health endpoints; remote exposure needs additional
production controls.

v5.0.1

Choose a tag to compare

@SAY-5 SAY-5 released this 27 Sep 23:46
ee2e37d

Two tests compared a regenerated artifact with the committed one exactly, and that does not hold across CPU architectures: regenerated from its seed on the x86-64 CI runner, tests/fixtures/replay_log.jsonl gets distance_km 8.5 on the row where the committed file records 8.499, and the distance_km standard deviation in artifacts/manifest.json comes out 4.990105 against the recorded 4.990104, while an arm64 Mac reproduces the fixture byte for byte and the statistics exactly. test_fixture_is_reproducible_from_the_seeded_dataset and test_manifest_carries_training_stats_for_every_input now compare through assert_reproduces in tests/conftest.py, which holds each number to a tolerance declared per field: 0.001 for distance_km and 0.0001 for traffic_index, one unit in the last place modelgate/model/data.py rounds them to, 0.000001 for every figure modelgate/model/stats.py rounds to six decimals, and 0.02 for the two eta fields derived from those inputs. Keys, field order, line count, integers, strings and booleans still have to match exactly, and a decimal with no tolerance declared for it or for a container above it is refused, so a new field cannot pass unchecked. tests/test_artifact_tolerance.py pins the comparison: a difference inside the tolerance passes in either direction, two units in the last place, a drifted integer, a flipped boolean, a missing key or a changed shape fail, and a number outside its tolerance is reported with the field, the difference and the tolerance.

The README, the assert_reproduces docstring and the test comments that explain this comparison now describe it as the artifacts and the code show it. They had the direction reversed, saying the committed fixture records 8.5 where the runner regenerates 8.499; they named torch's vectorised exp as the cause, which no artifact isolates; they called every tolerance one unit in the last place, which the two eta fields are not; and they said a number with no declared tolerance must match exactly, where the code refuses it outright. The test pinning that refusal is renamed test_a_decimal_without_a_declared_tolerance_is_refused_even_when_equal. tests/measure_eta_shift.py makes the basis of the 0.02 eta tolerance reproducible: moving distance_km and traffic_index by one unit in their last place together, in every sign combination over the 300 fixture records, shifts the reference eta by at most 0.009381 minutes and the prediction v1 serves by at most 0.002720. The model and its artifacts are unchanged, so artifacts/manifest.json still records the 5.0.0 build that trained them.

Two timing budgets are widened. test_healthy_canary_stays_active, which failed on the CI runner when its healthy v2 canary was rolled back, now raises latency_floor_ms to 250 ms: a latency rollback needs the candidate's p95 to be over twice the primary's and more than that floor above it, and test_slow_canary_rolls_back_on_latency still proves the rule by injecting a 4 ms delay into the candidate. The batching test's max_wait rises from 5 ms to 50 ms with its assertions unchanged.

The browser demo under web/ moves from Vite 5.4.21 to 7.3.6 and @vitejs/plugin-react from 4.7.0 to 5.2.0 and declares node ^20.19.0 || >=22.12.0; npm audit over web/package-lock.json reported esbuild (moderate, GHSA-67mh-4wv8-2f99) and Vite (high, three advisories including GHSA-fx2h-pf6j-xcff) at v5.0.0 and reports 0 vulnerabilities at this commit, where the demo type-checks, bundles and passes its 15 self-check assertions, and a new CI web job runs npm ci under the Node 22 that web/.nvmrc names, the type-check, the self check and the bundle on every push. 116 tests at this commit, passing on the x86-64 CI runner and on an arm64 Mac, and the load test with shadow, pre-warm, a 0.1 canary and a mid-run swap drops 0 of 1800 requests.

v5.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:59

Adds a sampled request log and an offline replay harness. With MODELGATE_REQUEST_LOG set, an exact deterministic share of accepted requests is written as JSON lines containing only the six validated numeric inputs, the serving version, and the answer; nothing from the HTTP layer reaches the file. modelgate eval replays such a log through the same validation, encoding, and padded forward pass the service uses and reports MAE, p95 error, bias, a calibration table with expected calibration error per version, the divergence between two versions, and how many logged answers the replay reproduced (0 mismatches on the committed 300-record fixture, where v1 scores MAE 2.94 and v2 2.01). The README now carries a Releases table summarizing 1.0.0 through 5.0.0. 104 tests.

v4.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:53

Adds micro-batching and a warm model pool. Concurrent /predict calls for the same model are queued FIFO on the event loop and run in one forward pass, bounded by a maximum batch size and wait; every forward pass is padded to a fixed row count so a request gets bit-identical output alone or in a full batch, which the swap test checks for all 2000 responses across a promote. Sparse traffic skips the wait, so single-row p50 stays at 1.7 ms. POST /admin/warm loads a candidate into a spare pool slot ahead of time and the following promote reports prewarmed with 0 s of load; role holders are never evicted. 93 tests, load test with pre-warm, canary, and mid-run swap passes with 0 drops.

v3.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:39

Adds feature drift monitoring. The trainer now writes per-input reference statistics (mean, std, quantiles, decile edges, category frequencies) into manifest.json, and the serving layer scores a rolling window of accepted inputs against them with a population stability index per feature, plus an unknown-category rate fed by validation rejections. GET /admin/drift returns scores, statuses, and live versus training statistics; the same numbers are exported as modelgate_feature_drift gauges. Training-like traffic scores about 0.01 per feature and a 3x shift in trip distance scores 2.3 on distance_km alone. 76 tests.

v2.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:34

Adds canary routing. POST /admin/canary sends a fixed share of live traffic to a candidate version, using a deterministic weighted split so a 5% weight really is every 20th request. The router keeps a rolling window of error rate and p95 inference latency per version and rolls the canary back on its own when the candidate exceeds the primary by the configured margin; a failing candidate answers from the primary for that request, so rollback happens with 0 dropped requests. 67 tests, load test with canary plus mid-run swap passes with 0 drops.

v1.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:28

First stable release of the ETA serving service. It serves POST /predict with strict input validation, runs an optional shadow version on every request and reports divergence, and swaps model versions atomically so that in-flight requests finish on the model they started with. Prometheus metrics and a provisioned Grafana dashboard cover traffic, latency, rejections, shadow divergence, swaps, and drops; the load test promotes v2 mid-run with 0 dropped requests.