Skip to content

Harness Evals 0.1.0

Choose a tag to compare

@Dhi13man Dhi13man released this 12 Jul 07:52
1b516cb

Harness Evals 0.1.0 is the first public release of a generic evaluator for software engineering and testing instruction bundles across coding harnesses.

Highlights

  • Seventeen calibrated implementation and testing cases with objective, case-specific verifiers and known-good, known-bad, and adversarial variants.
  • Isolated provider execution, Git-bound source snapshots, blinded AB/BA comparison, bounded spend accounting, and externally sealed holdout-plan support.
  • Neutral engineering and testing reference bundles with no dependency on private skill names or private repository history.
  • Installable Python package with harness-evals, python -m harness_evals, and harness-evals-prepare-holdout entry points.
  • MIT licensing, contribution and security policies, CodeQL, OpenSSF Scorecard, Dependabot, protected-main rules, and reproducible CI release gates.

Scope

This is an alpha release for expert use on Linux with Python 3.11 or newer, a working systemd --user manager, and unprivileged user/mount namespaces. The public corpus contains train and validation cases, not a private holdout. No live comparator certification is included, and this release makes no claim that one harness or instruction bundle is superior.

See the 0.1.0 changelog for the complete release surface.