Skip to content

Bug Receipt v1.2.0 — reproducible benchmarks

Choose a tag to compare

@lMysticl lMysticl released this 12 Aug 00:05
· 18 commits to main since this release

Measured evidence

  • Deterministic receipt invariants: 20/20 passed.
  • Fresh Codex routing probes: 3/3 selected Bug Receipt.
  • Paired skills-ON quality: 12/12 assertions passed.
  • Paired skills-OFF quality: 12/12 assertions also passed — this bounded sample does not prove a quality uplift.
  • Raw telemetry, fingerprints, exact prompts, and the quarantined failed holdout are published in the repository.

Also included

  • A more precise BLOCKED evidence-package rule for cross-system failures.
  • A public benchmark section on the project site and llms.txt.
  • Reproducible validator benchmark via npm run benchmark:validator.

Full report: https://github.com/lMysticl/bug-receipt/blob/main/benchmarks/RESULTS.md