Skip to content

Harness Evals 0.2.0

Choose a tag to compare

@Dhi13man Dhi13man released this 12 Jul 09:15
70ddc67

Harness Evals 0.2.0 makes the evaluator boundary match its intended scope: reproducible A/B evaluation of user-defined skills and instruction bundles through agent harnesses.

Changed

  • Treats the bundled engineering and testing tracks as the first reference corpus, not the product boundary.
  • Derives holdout skill counts and aggregate cells from each selected suite instead of two built-in identifiers.
  • Uses domain-neutral isolated task instructions so other skill domains are not framed as software-engineering work.
  • Documents Claude and Codex as explicitly registered built-in adapters; Codex remains optional and diagnostic.
  • Refreshes release locks, package metadata, issue forms, citation metadata, and supported-version policy for 0.2.0.
  • Makes test fixture Git operations synchronous to eliminate a Python 3.11 cleanup race.

The bundled blinded comparator remains calibrated for software-change evidence. General comparator profiles, provider capability registration, alternate bundle sources, and installed resource loading are tracked in issue #4 rather than being overclaimed in this release.

Published artifacts come from the successful protected-main CI run for commit 70ddc67bcf52ad1c4262dd7166ce6c28c7ab9e9c.