Skip to content

v0.5.0

@alexfilatov alexfilatov tagged this 16 May 08:23
240-trial benchmark against @gymbile/wpl-ai ^1.13.0 across 4 OpenAI models
× 15 scenarios × 2 lanes × 2 phases. Total reproduce cost: $37.27.

Headlines:
  - Lane A unsafe: 43/120 (36%); 207 violations
  - Lane B unsafe:  6/120 (5%);  28 violations
  - Reduction: 86% on both metrics
  - Lane B served: 109/120 (91%), 64/120 complete (≥10 wk)
  - Multi-turn drift: Lane A 25/60 (42%), Lane B 0/60

Changes from v0.4 (full disclosure in docs/DIFF_v0.4_to_v0.5.md):
  - Repaired 10 dead scoring blacklist entries (matched nothing before)
  - Repaired Lane A extractor truncation (silently zeroed 27 of 120
    trials in v0.4)
  - Persisted extractor raw output for offline-recoverable parse
    failures
  - Bumped @gymbile/wpl-ai 1.12 -> 1.13 (stricter compiler)
  - Updated all 4 publication docs + audit + diff + roadmap
  - Four hero charts in docs/charts/

Companion private repo: github.com/gymbile/gymbile-internal
(operational content not in this public repo)
Assets 2
Loading