v0.2.0 — Verified local-first portfolio release
Correction: upgrade to v0.2.1. Version 0.2.0 contains a stale-approval defect in enforcement of version-wide blocks. The corrective release fixes it; its later evidence addendum records completion under the owner-approved AI-authored benchmark scope, without claiming independent human validation. The historical packages, evidence assets and tag below remain unchanged.
This release turns the evaluation harness into a runnable, documented local-first portfolio project with an authenticated shared-server option.
- Complete trace verification, fenced execution, atomic checkpoints, comparable baselines and authenticated tenant isolation.
- Passing Linux Python 3.12/3.13/3.14 and Windows Python 3.12 CI; 281 tests pass per matrix job, with eight additional PostgreSQL contracts passing in the database job.
- Verified installed wheel, Docker behavior, persistent Compose/MCP restart and private HTTPS through Caddy with explicit CA trust.
- Static 150-second walkthrough, engineering case study, operational checks and real OpenTelemetry Collector evidence.
Interactive demo | Original release verification | Release CI
The first live Sonnet 5 experiment is preserved: 21 paid requests, $0.10417 estimated token cost, 1/10 cases passed, and a review-required gate. These results expose fixture/latency limits and are not a general model-quality score. No further paid runs were needed.
Alexander Hines approved all 33 scorer challenge rows and 35 labels without corrections on September 27, 2026. Nine semantic disagreements remain documented; this release does not claim independently human-validated scoring accuracy or a publicly hosted evaluation backend.
Assets include the wheel, source archive, SHA256SUMS and a verification ZIP containing test reports, coverage, source provenance, GitHub run records and bounded live/operational evidence. SHA256SUMS covers the original wheel, source archive and verification ZIP; those four assets remain unchanged.
Human-review addendum: review sheet and successful review CI. The separate human-review-addendum-0.2.0.zip and HUMAN-REVIEW-SHA256SUMS bind the reviewed corpus, unchanged scorer results, completed review sheet and test reports to commit 517f96c. All six CI jobs pass: 282 tests per OS/Python matrix job and eight additional PostgreSQL tests. The corpus remains development-exposed, and all nine disagreements are preserved. The existing label review is complete. Independent human evidence was pending at this historical snapshot. The owner subsequently approved AI-authored characterization instead; see the current v0.2.1 evidence addendum. Independent human validation is not claimed.