Skip to content

v0.1.0 — Same model, different wrapper

Choose a tag to compare

@minghinmatthewlam minghinmatthewlam released this 03 Jul 13:45
· 851 commits to main since this release

First release: the complete M3→M4 benchmark arc.

Findings: harness efficiency separates up to 8× in tokens where correctness saturates; 3 of 4 open models (GLM-5.2, DeepSeek V4 Flash, Kimi K2.7) reach frontier parity on our tasks — the entire 72-run open-model matrix cost ~$1.02.

In the box: 5 harness adapters (codex, pi, opencode, cursor, devin), open-model support via first-party APIs, validated partial-credit tasks + an imported Exercism tier, Docker isolation mode, Wilson-CI statistics, transcript persistence with PII scrubbing, and 4 committed datasets.

📖 Start with WRITEUP.md · 📊 Live results · 🔁 Reproduce for ~$1: see README.