Skip to content

v0.6.0

Latest

Choose a tag to compare

@github-actions github-actions released this 06 Jul 16:27
bcce39a

Highlights

The first feature release since v0.5.1. Two themes: a harder task tier that un-saturates the frontier, and the tooling to benchmark it affordably.

Frontier-hard task tier — net-new-algorithm tasks from real post-cutoff merged PRs (networkx Leiden, sqlglot canonicalize, three PennyLane resource-estimation templates), each individually screened to defeat low-effort frontier models, with grading calibrated to be fair-hard (difficult, not under-specified).

Cost-efficient reporting — a new-model report becomes one column instead of a full-matrix re-run, because grading is deterministic and every run records the task hash it was scored against:

  • --max-run-cost — per-run cost ceiling; stop and grade the partial result
  • compare — assemble the model × effort matrix from cached runs (baselines never re-run)
  • regrade — re-score existing runs against the current tests for $0
  • --only-missing — resume a suite run, filling only the uncached cells

Net effect: a new-model report drops from ~$100 to ~$10. See docs/QUICKSTART.md for the workflow.


Install

git clone https://github.com/morganlinton/VulcanBench.git
cd VulcanBench
make setup
source .venv/bin/activate
vulcanbench --help

Build the sandbox image once:

make sandbox-image

See CHANGELOG.md for full notes.

What's Changed

  • v2: realistic hard tier - 10 real-PR tasks, both frontier models in the 70-90% band, $7.36/run by @morganlinton in #2
  • Anthropic provider: refusal guard and output headroom for Fable 5 by @morganlinton in #3
  • Add published benchmark results to docs/results by @morganlinton in #4
  • Add 5 hard multi-site v2 task candidates (anyio, pytest, urllib3, celery, redis) by @morganlinton in #5
  • Discount OpenAI cached prompt tokens in cost accounting by @morganlinton in #7
  • Add 11 post-cutoff OSS algorithm tasks: veryhard tier reaches 5 confirmed frontier-hard by @morganlinton in #8
  • Calibrate the veryhard tier: fair-hard instead of under-specified by @morganlinton in #9
  • Add per-run cost cap and zero-cost regrade command by @morganlinton in #10
  • Add compare: assemble model×effort matrix from cached runs (freeze baselines) by @morganlinton in #11
  • Add --only-missing: resume a suite run, filling only uncached cells by @morganlinton in #12
  • docs: note the runs-dir consolidation caveat for --only-missing and compare by @morganlinton in #13
  • docs: document cost-efficient commands in the README by @morganlinton in #14
  • Release v0.6.0: frontier-hard tier + cost-efficient reporting by @morganlinton in #15

New Contributors

Full Changelog: v0.1.0...v0.6.0