Skip to content

v0.25.1 — docs: EV-007c main run published (learning events 3/3; sealed line not met, outcome row 8)

Choose a tag to compare

@shojikumaru shojikumaru released this 05 Sep 19:35
· 33 commits to main since this release
109154f

Docs-only release: the EV-007c main run is published under docs/benchmark.md#ev-007 (EN + JA), and the four README EV-007 paragraphs are aligned.

What it says. EV-007c re-ran the sealed EV-007b series on v0.25.0, the pin that carries the #274 citation-contract fix. It is the first run of the family with learning events — 3 loop turns, 3 successful reviews (0 fail-closes), 6 promotions, 2 approved rules — so the sealed pass/fail line was evaluable for the first time. Sealed outcome row 8 — fail: the pre-registered line was not met (learning-arm repeat-mistake rate 0.667 vs 0.727 in the no-learning arm; DiD +0.050; attribution rule met; no-harm guard intact; cost 79,070,996 = 93.0 % of the 85M cap; no post-seal correction). The one T2 class that moved after the rule was approved (in-file supersession) was solved in round 3 and half-lost in the final block; the erratum-file classes never moved for any arm and were never written down for the loop to see. Measured once, n = 1, one bit per class — a fail of that line, not a verdict on the loop.

Evidence. PR #276 (completion record L1-7; 5-seat blind panel — Gemini / Muse / GLM / Codex / Grok — cumulative GO; CI green incl. test-macos 26m36s: https://github.com/caty-ai/caty-agent-harness/actions/runs/33985419350). Home issue #264. Previous release: v0.25.0 (#274 / PR #275).