Releases: SAY-5/panelist
Release list
v6.0.0
Two changes break callers of 5.0.0, which is why this is a major release. The dataset export moved from GET /deliveries/export to POST /deliveries, since every call stores a new version; it answers 201, and 6.0.0 answers 404 at the old path. Experts no longer carry hourly_rate_cents: POST /experts ignores it and no expert response returns it. The column itself stays in the table, unread and unwritten, because 5.0.0 selects and inserts it and the Terraform deployment runs alembic upgrade head in each new task while 5.0.0 tasks keep serving; a 7.0.0 migration will drop it. Against a database upgraded to 6.0.0, 5.0.0 answered GET /experts/me with 200, POST /tasks/next with 204, GET /experts/{id} with 200 and POST /experts with 201. tests/test_migrations.py upgrades from the 5.0.0 schema and fails if any column it had is gone.
Service errors derive from one ServiceError that a single FastAPI handler maps to its status, so an approval that finds no rate for the expert's tier and the task's type answers 422 on a review or an adjudication, where 5.0.0 answered 500; tests/test_payouts.py checks that the failed review leaves the grade unreviewed and unpaid. A single-grader task whose grade a reviewer rejects returns to the queue for a different expert instead of ending in rejected. Rejecting the grade an agreed consensus round selected picks the delivered grade again among the approved ones, where 5.0.0's export skipped such a task although it was approved. The scheduler tick reports tasks rejected twice and approved tasks whose selected grade is still unreviewed. GET /deliveries/{version}/verify reads the stored object back from S3 or disk and recomputes its sha256, row count and size, and listing and verifying need the new deliveries:read scope, which reviewers, senior reviewers and admins hold. DELETE /admin/api-keys/{id} revokes a key from the next request on, and global agreement is computed in one query instead of one per task.
Attention checks are no longer served on a countable cadence. 5.0.0 made every round(1/f)-th serve prefer a golden task while its routing code said the expert cannot tell the difference, so an expert who counted serves knew which serves could carry a check. 6.0.0 decides each serve with sha256 over ATTENTION_KEY, the expert id and the serve number, compared with ATTENTION_FRACTION, which holds the configured share without a pattern to count. That hides the schedule only from someone without the key, which defaults to panelist in the code and to change-me in Terraform, so each deployment has to set its own; Terraform keeps it in Secrets Manager and injects it into the task as ATTENTION_KEY. tests/test_attention.py drives an expert who is careful exactly on the serves the old cadence predicts and asserts the guard still pauses them.
The engine no longer lets psycopg prepare statements on the server (prepare_threshold=None in panelist/db.py). psycopg prepares a statement once a connection has run it five times since its last rollback, and a prepared statement fails with cached plan must not change result type once a migration changes the type of a column it returns. A 5.0.0 uvicorn --workers 4 server left running across alembic downgrade base and upgrade head, which recreate the role enum the API key lookup returns, answered 2 of 120 PUT /rate-cards requests after the reset with 500; 6.0.0, run the same way, answered all 120 with 204. tests/test_db.py runs the key lookup ten times on one pooled connection across that reset. In the claim benchmark, four alternating runs each of bc32218, the last commit that prepared statements, and e394ea6 overlap on every figure (docs/prepared-statements-2026-09-27-solo.json, docs/prepared-statements-2026-09-27-contended.json).
sim/bench.py measures the claim path with one claimant and then with all of them, and sim/bench_session.py runs alternating rounds on a single in-process worker, which shares its interpreter with the claimant threads, and on uvicorn --workers 4, resetting the schema before every run. Two sessions are committed, docs/bench-2026-09-27.json and docs/bench-2026-09-27-2.json, each run making 50 claims with one claimant and then 520 with 40. The sessions disagree on absolute figures: every in-process round of the second had a higher p50 and a lower throughput than every in-process round of the first, its contended p50 18 to 46 percent above the first session's highest, at loads inside the first session's range. What held in both is that at 40 claimants every four-worker p50 was below every one-worker p50, by 2.1 to 3.4 times within a pair of rounds, four workers claimed 1.6 to 2.5 times as fast, and no task was handed to two claimants. The rows do not say why: the four-worker server runs the claim path in four interpreters, with FOR UPDATE SKIP LOCKED keeping their claims apart, while the in-process one shares its interpreter with 40 claimant threads, so they cannot separate process placement from SKIP LOCKED. The contended p95 and the single-claimant figures do not tell the servers apart.
The documentation quotes recorded runs. The README demo block is one run recorded in docs/demo-2026-09-26.json, and make demo-check rebuilds its world from seed 7 and compares fingerprints; in 5.0.0 each expert's grading noise came from Python's per-process salted hash of the name, so seed 7 did not fix it. The README API and configuration tables, which in 5.0.0 left out seven of 41 routes and four settings, now match the code, and tests/test_docs.py fails when they, a quoted benchmark figure or the quoted test count drift. The browser port under web/ is a TypeScript port of the 1.0.0 service layer, which web/src/sim/port.ts names; the 5.0.0 README said it printed the same summary block as the service and the page labelled its counters measured results, and the page now labels them a simulated run. tests/test_port_conformance.py records one fixed scenario run through the service and npm run selfcheck replays it through the port, and a web CI job runs the typecheck, the selfcheck, which passed 123 of 123 assertions at this commit, the production bundle and a gzip budget. esbuild moves from 0.24.2 to 0.25.12, clearing GHSA-67mh-4wv8-2f99, and .python-version pins CPython 3.12, the version CI and the image use. 74 tests pass at this commit, against PostgreSQL 16 in CI.
v5.0.0
GET /ops/overview answers the questions an operator asks first, in one read: queue depth by required tag, task counts by status, how many assigned tasks have already run past their lease, which experts are paused and what calibration score they were carrying, how many tasks sit in the adjudication queue, what has not yet been swept into a payout statement, and which delivery version was published last. uv run panelist tick is the scheduler pass meant for cron: it reclaims expired leases, recomputes every expert's rolling calibration score so a score never goes stale because nobody happened to review that expert, and prints a JSON report whose reminders name the experts still short of the minimum sample count, the payouts sitting outside a statement, and the adjudications still waiting for a decision. GET /ops/audit.csv exports the append-only audit trail as CSV for admins, filtered by action and start time. New Prometheus gauges cover the adjudication backlog, the paused expert count and the expert count per tier, alongside a counter for automatic promotions and demotions. make demo now settles its disputed tasks through a senior reviewer and ends with the tick report and the overview, and the README shows the real output of that run. The suite is 56 tests against PostgreSQL 16, including exact overview counts on a seeded five task fixture.
v4.0.0
A task can now ask for k graders, and when the last required grade lands Panelist compares the weighted scores instead of leaving the decision to whoever reviews first. If the spread stays inside CONSENSUS_TOLERANCE the round is recorded as agreed and the grade closest to the mean becomes the one that will be delivered. If the spread is wider the task moves to the new adjudication status and appears at GET /adjudications with every grade, the spread and the tolerance that was in force, and ordinary reviews on it are refused with 409 until it is decided. POST /adjudications/{task_id} takes the delivered grade and a reason, and needs the new adjudications:write scope, which only the new senior_reviewer role and admins hold; the chosen grade is approved and delivered, and every other grade is recorded as outvoted. Outvoted graders are paid by CONSENSUS_OUTVOTED_PAYOUT, either the full card rate, a CONSENSUS_OUTVOTED_RATE fraction of it, or nothing at all. The delivery export now emits one row per consensus task, the delivered grade, carrying the round status, the grader count and the spread; the suite is 52 tests against PostgreSQL 16.
v3.0.0
Every reviewer decision and every golden check result now counts as an agreement signal for the expert who produced the grade. Panelist keeps a rolling window of those signals, stores the agreement rate and sample count on the expert, and exposes both at GET /experts/{id}/calibration together with the full tier change history. Once enough signals exist, a rate at or above the promote edge moves the expert up one tier and a rate at or below the demote edge moves them down one; rates in the band between the two edges leave the tier untouched, so a single decision either way cannot make an expert flap across a boundary. Routing, direct claims and payout rate lookups all read the live tier, so a demoted expert stops being offered tasks above their new tier on the very next claim and a direct claim on one returns 403. Every move writes a tier_changes row and an audit event. Four new API level tests cover promotion after agreeing grades, demotion after a disagreement streak, the no flapping band and tier aware routing; the suite is 48 tests against PostgreSQL 16.
v2.0.0
Panelist 2.0.0 makes rubric versions immutable. Publishing a new version through POST /rubrics/{id}/versions creates the next version, moves every untouched queued task onto it, and reports how many open tasks still reference the previous version. A task pins its rubric version at claim time, so an expert who was shown version 1 keeps grading version 1 even when version 2 is published during the lease, and the stored grade references the pinned version. A grade that names a superseded version is rejected with a 409 that states both versions, new tasks cannot target a superseded version, and only the current version can be published from. Criterion analytics group by rubric version and delivery rows carry the version each grade was produced against. The suite grows to 44 tests, all exercised through the HTTP API against PostgreSQL.
v1.0.0
Panelist 1.0.0 is the baseline release of the expert grading and data delivery platform. Experts pull tasks from a queue routed by expertise tag with row level locking, leases and reclaim, then grade model outputs against weighted rubrics. Hidden attention checks track a rolling pass rate that pauses careless experts and withholds their payouts. Approved grades create payouts at tier and task type rates, close into period statements with CSV export, and are delivered as versioned, checksummed JSONL datasets to S3 or disk. Terraform provisions ECS Fargate, RDS PostgreSQL and S3, and the test suite of 41 tests runs against PostgreSQL through Testcontainers.