Agent Preflight v0.3.0 upgrades the skill pack from self-reported routing checks to a reproducible paired, reviewer-blinded evaluation workflow.
Highlights
- Added paired control-versus-preflight trials with identical requests and opaque IDs.
- Added anonymous A/B review queues, separate blind keys, and SHA-256 artifact binding.
- Added automatic unblinding with paired score deltas, bootstrap 95% confidence intervals, two-sided sign tests, win rates, hard-failure comparison, and complete-pair token/latency/tool-call metrics.
- Expanded the behavior corpus from 39 to 60 adversarial cases: 32 English and 28 Chinese.
- Improved prompt-injection handling, contradiction detection, scoped approvals, local re-entry, and high-impact gate questioning.
- Reduced visible preflight ceremony for trivial and inspectable work.
Validation
- 6 skills pass the official skill-format validator.
- 60-case corpus validation passes.
- 26 unit and integration tests pass on Python 3.8 and 3.12.
- Full-corpus preparation produces 120 isolated trials across 60 pairs.
- PowerShell and Git Bash installation paths were verified.
Evidence note
The evaluation machinery is validated, but universal product-effect claims still require fresh, representative runs with at least 30 complete pairs and independent blinded review.
See the evaluation guide and changelog.