Skip to content

Agent Preflight v0.3.0 - Blind paired evaluation

Latest

Choose a tag to compare

@wde123sadw wde123sadw released this 16 Jul 06:09

Agent Preflight v0.3.0 upgrades the skill pack from self-reported routing checks to a reproducible paired, reviewer-blinded evaluation workflow.

Highlights

  • Added paired control-versus-preflight trials with identical requests and opaque IDs.
  • Added anonymous A/B review queues, separate blind keys, and SHA-256 artifact binding.
  • Added automatic unblinding with paired score deltas, bootstrap 95% confidence intervals, two-sided sign tests, win rates, hard-failure comparison, and complete-pair token/latency/tool-call metrics.
  • Expanded the behavior corpus from 39 to 60 adversarial cases: 32 English and 28 Chinese.
  • Improved prompt-injection handling, contradiction detection, scoped approvals, local re-entry, and high-impact gate questioning.
  • Reduced visible preflight ceremony for trivial and inspectable work.

Validation

  • 6 skills pass the official skill-format validator.
  • 60-case corpus validation passes.
  • 26 unit and integration tests pass on Python 3.8 and 3.12.
  • Full-corpus preparation produces 120 isolated trials across 60 pairs.
  • PowerShell and Git Bash installation paths were verified.

Evidence note

The evaluation machinery is validated, but universal product-effect claims still require fresh, representative runs with at least 30 complete pairs and independent blinded review.

See the evaluation guide and changelog.