Harness Evals 0.1.0
Harness Evals 0.1.0 is the first public release of a generic evaluator for software engineering and testing instruction bundles across coding harnesses.
Highlights
- Seventeen calibrated implementation and testing cases with objective, case-specific verifiers and known-good, known-bad, and adversarial variants.
- Isolated provider execution, Git-bound source snapshots, blinded AB/BA comparison, bounded spend accounting, and externally sealed holdout-plan support.
- Neutral
engineeringandtestingreference bundles with no dependency on private skill names or private repository history. - Installable Python package with
harness-evals,python -m harness_evals, andharness-evals-prepare-holdoutentry points. - MIT licensing, contribution and security policies, CodeQL, OpenSSF Scorecard, Dependabot, protected-main rules, and reproducible CI release gates.
Scope
This is an alpha release for expert use on Linux with Python 3.11 or newer, a working systemd --user manager, and unprivileged user/mount namespaces. The public corpus contains train and validation cases, not a private holdout. No live comparator certification is included, and this release makes no claim that one harness or instruction bundle is superior.
See the 0.1.0 changelog for the complete release surface.