Bug Receipt v1.2.0 — reproducible benchmarks
Measured evidence
- Deterministic receipt invariants: 20/20 passed.
- Fresh Codex routing probes: 3/3 selected Bug Receipt.
- Paired skills-ON quality: 12/12 assertions passed.
- Paired skills-OFF quality: 12/12 assertions also passed — this bounded sample does not prove a quality uplift.
- Raw telemetry, fingerprints, exact prompts, and the quarantined failed holdout are published in the repository.
Also included
- A more precise
BLOCKEDevidence-package rule for cross-system failures. - A public benchmark section on the project site and
llms.txt. - Reproducible validator benchmark via
npm run benchmark:validator.
Full report: https://github.com/lMysticl/bug-receipt/blob/main/benchmarks/RESULTS.md