Harness Evals 0.2.0
Harness Evals 0.2.0 makes the evaluator boundary match its intended scope: reproducible A/B evaluation of user-defined skills and instruction bundles through agent harnesses.
Changed
- Treats the bundled engineering and testing tracks as the first reference corpus, not the product boundary.
- Derives holdout skill counts and aggregate cells from each selected suite instead of two built-in identifiers.
- Uses domain-neutral isolated task instructions so other skill domains are not framed as software-engineering work.
- Documents Claude and Codex as explicitly registered built-in adapters; Codex remains optional and diagnostic.
- Refreshes release locks, package metadata, issue forms, citation metadata, and supported-version policy for 0.2.0.
- Makes test fixture Git operations synchronous to eliminate a Python 3.11 cleanup race.
The bundled blinded comparator remains calibrated for software-change evidence. General comparator profiles, provider capability registration, alternate bundle sources, and installed resource loading are tracked in issue #4 rather than being overclaimed in this release.
Published artifacts come from the successful protected-main CI run for commit 70ddc67bcf52ad1c4262dd7166ce6c28c7ab9e9c.