Skip to content

Verity Benchmark v0.2

Latest

Choose a tag to compare

@Th0rgal Th0rgal released this 15 Aug 07:26
c5a2344

Detailed verifier-backed artifacts for benchmark version 0.2.

This release contains the frozen 240-task manifest, the deterministic STRAT-50 panel, a machine-readable leaderboard, per-cohort run bundles, and SHA-256 checksums. Every scored cohort has 50 distinct terminal verifier verdicts on the same panel with the p4_normal budget of 16 attempts and 120 tool calls. Infrastructure-invalid attempts are preserved where available and excluded from scores.

The primary table selects the highest-scoring completed protocol for each model family. Lower-scoring protocol variants remain separate audit artifacts. Kimi ordinary Chat scored 14/50 versus 12/50 with reasoning_content replay. Muse Responses scored 10/50 versus 1/50 with Chat.

STRAT-50 reports the direct observed solve rate on the frozen 50-task panel. It is not an unweighted estimate of FULL-240 performance. Benchmark v0.2 uses Lean 4.24.0 and is retained as a historical reproducibility snapshot. New campaigns should use benchmark v0.3 with Lean 4.31.