240-trial benchmark against @gymbile/wpl-ai ^1.13.0 across 4 OpenAI models
× 15 scenarios × 2 lanes × 2 phases. Total reproduce cost: $37.27.
Headlines:
- Lane A unsafe: 43/120 (36%); 207 violations
- Lane B unsafe: 6/120 (5%); 28 violations
- Reduction: 86% on both metrics
- Lane B served: 109/120 (91%), 64/120 complete (≥10 wk)
- Multi-turn drift: Lane A 25/60 (42%), Lane B 0/60
Changes from v0.4 (full disclosure in docs/DIFF_v0.4_to_v0.5.md):
- Repaired 10 dead scoring blacklist entries (matched nothing before)
- Repaired Lane A extractor truncation (silently zeroed 27 of 120
trials in v0.4)
- Persisted extractor raw output for offline-recoverable parse
failures
- Bumped @gymbile/wpl-ai 1.12 -> 1.13 (stricter compiler)
- Updated all 4 publication docs + audit + diff + roadmap
- Four hero charts in docs/charts/
Companion private repo: github.com/gymbile/gymbile-internal
(operational content not in this public repo)