Skip to content

v6.0.0

Latest

Choose a tag to compare

@SAY-5 SAY-5 released this 28 Sep 18:25
4bf3bc0

Two changes break callers of 5.0.0, which is why this is a major release. The dataset export moved from GET /deliveries/export to POST /deliveries, since every call stores a new version; it answers 201, and 6.0.0 answers 404 at the old path. Experts no longer carry hourly_rate_cents: POST /experts ignores it and no expert response returns it. The column itself stays in the table, unread and unwritten, because 5.0.0 selects and inserts it and the Terraform deployment runs alembic upgrade head in each new task while 5.0.0 tasks keep serving; a 7.0.0 migration will drop it. Against a database upgraded to 6.0.0, 5.0.0 answered GET /experts/me with 200, POST /tasks/next with 204, GET /experts/{id} with 200 and POST /experts with 201. tests/test_migrations.py upgrades from the 5.0.0 schema and fails if any column it had is gone.

Service errors derive from one ServiceError that a single FastAPI handler maps to its status, so an approval that finds no rate for the expert's tier and the task's type answers 422 on a review or an adjudication, where 5.0.0 answered 500; tests/test_payouts.py checks that the failed review leaves the grade unreviewed and unpaid. A single-grader task whose grade a reviewer rejects returns to the queue for a different expert instead of ending in rejected. Rejecting the grade an agreed consensus round selected picks the delivered grade again among the approved ones, where 5.0.0's export skipped such a task although it was approved. The scheduler tick reports tasks rejected twice and approved tasks whose selected grade is still unreviewed. GET /deliveries/{version}/verify reads the stored object back from S3 or disk and recomputes its sha256, row count and size, and listing and verifying need the new deliveries:read scope, which reviewers, senior reviewers and admins hold. DELETE /admin/api-keys/{id} revokes a key from the next request on, and global agreement is computed in one query instead of one per task.

Attention checks are no longer served on a countable cadence. 5.0.0 made every round(1/f)-th serve prefer a golden task while its routing code said the expert cannot tell the difference, so an expert who counted serves knew which serves could carry a check. 6.0.0 decides each serve with sha256 over ATTENTION_KEY, the expert id and the serve number, compared with ATTENTION_FRACTION, which holds the configured share without a pattern to count. That hides the schedule only from someone without the key, which defaults to panelist in the code and to change-me in Terraform, so each deployment has to set its own; Terraform keeps it in Secrets Manager and injects it into the task as ATTENTION_KEY. tests/test_attention.py drives an expert who is careful exactly on the serves the old cadence predicts and asserts the guard still pauses them.

The engine no longer lets psycopg prepare statements on the server (prepare_threshold=None in panelist/db.py). psycopg prepares a statement once a connection has run it five times since its last rollback, and a prepared statement fails with cached plan must not change result type once a migration changes the type of a column it returns. A 5.0.0 uvicorn --workers 4 server left running across alembic downgrade base and upgrade head, which recreate the role enum the API key lookup returns, answered 2 of 120 PUT /rate-cards requests after the reset with 500; 6.0.0, run the same way, answered all 120 with 204. tests/test_db.py runs the key lookup ten times on one pooled connection across that reset. In the claim benchmark, four alternating runs each of bc32218, the last commit that prepared statements, and e394ea6 overlap on every figure (docs/prepared-statements-2026-09-27-solo.json, docs/prepared-statements-2026-09-27-contended.json).

sim/bench.py measures the claim path with one claimant and then with all of them, and sim/bench_session.py runs alternating rounds on a single in-process worker, which shares its interpreter with the claimant threads, and on uvicorn --workers 4, resetting the schema before every run. Two sessions are committed, docs/bench-2026-09-27.json and docs/bench-2026-09-27-2.json, each run making 50 claims with one claimant and then 520 with 40. The sessions disagree on absolute figures: every in-process round of the second had a higher p50 and a lower throughput than every in-process round of the first, its contended p50 18 to 46 percent above the first session's highest, at loads inside the first session's range. What held in both is that at 40 claimants every four-worker p50 was below every one-worker p50, by 2.1 to 3.4 times within a pair of rounds, four workers claimed 1.6 to 2.5 times as fast, and no task was handed to two claimants. The rows do not say why: the four-worker server runs the claim path in four interpreters, with FOR UPDATE SKIP LOCKED keeping their claims apart, while the in-process one shares its interpreter with 40 claimant threads, so they cannot separate process placement from SKIP LOCKED. The contended p95 and the single-claimant figures do not tell the servers apart.

The documentation quotes recorded runs. The README demo block is one run recorded in docs/demo-2026-09-26.json, and make demo-check rebuilds its world from seed 7 and compares fingerprints; in 5.0.0 each expert's grading noise came from Python's per-process salted hash of the name, so seed 7 did not fix it. The README API and configuration tables, which in 5.0.0 left out seven of 41 routes and four settings, now match the code, and tests/test_docs.py fails when they, a quoted benchmark figure or the quoted test count drift. The browser port under web/ is a TypeScript port of the 1.0.0 service layer, which web/src/sim/port.ts names; the 5.0.0 README said it printed the same summary block as the service and the page labelled its counters measured results, and the page now labels them a simulated run. tests/test_port_conformance.py records one fixed scenario run through the service and npm run selfcheck replays it through the port, and a web CI job runs the typecheck, the selfcheck, which passed 123 of 123 assertions at this commit, the production bundle and a gzip budget. esbuild moves from 0.24.2 to 0.25.12, clearing GHSA-67mh-4wv8-2f99, and .python-version pins CPython 3.12, the version CI and the image use. 74 tests pass at this commit, against PostgreSQL 16 in CI.