v2.5.25 β benchmark v3 quality-first + evidence docs overhaul
What's new in v2.5.25
Benchmark v3 β quality-first (bukan cost-first).
- Defect density (anti-pattern findings per 100 LOC added) jadi metrik kualitas utama
- 2 task kompleks baru: multi-file refactor + 2-layer auth security bug
- Hasil terukur: no-rules agents GAGAL refactor (0/2), matcha satu-satunya arm yang selesai (1/1) β service layer diekstrak, semua test hijau
- Rework-loop counterfactual: 1 redo (~590K tokens) lebih besar dari seluruh premium matcha (+248K)
Docs overhaul β headline-first.
- Verdict box di atas, cost table dibingkai ulang sebagai 'invoice for verification'
- Section benchmark utuh: what-is / where-strong / which-feature-when / Run It Yourself
Standing-context slim.
- core.md 18.6K->15.9K chars + lazy-load + matcha-lite A/B arm (20 cells)
593 tests pass. Harness: node benchmark/live-bench.js --all --n 5