Bug Receipt v1.3.1 — verified, cross-platform audit gate
Bug Receipt v1.3.1 publishes the verified audit-gate benchmark and fixes cross-platform provenance verification.
Measured result:
- Skills ON: 18/20 (90%)
- Skills OFF: 7/20 (35%)
- Lift: +55 percentage points; exact paired McNemar p = 0.003418
- Complete receipts: 4/4; routing: 4/4; evidence-safety regressions: 0
- Independent corroborating run: 19/20 vs 7/20 (+60 points), p = 0.000488
The report SHA is now calculated over canonical LF text, so npm run benchmark:audit-gate is reproducible across Windows and Linux checkouts.
Scope: four frozen supplied-evidence closeout cases on gpt-5.6-sol/xhigh. This is a bounded result, not a universal ranking.