From the August 2026 harness audit (policy decision needed before the #139 v2 freeze).
Symptom
The response contract requires a short explanation per output, but the requirement is not enforced at scoring time: blank-explanation answers score normally.
Evidence (verified)
- 61 blank-explanation rows accepted on the current board: grok-4.3 = 56, kimi-k2.6 = 4, haiku-4.5 = 1.
- If blanks scored as misses, grok-4.3 would drop 2.574 pts (~7 ranks). No other model moves materially.
Decision for v2
Two coherent options:
- Score on numbers only (recommended): keep the explanation request in the prompt, publish explanation-compliance as a per-model column, but don't gate the accuracy score on prose compliance. Scoring prose mixes constructs into a numeric-accuracy headline.
- Enforce as miss, stated up front alongside the parse-failure policy (like Sonnet 5's 56/1,984 unparsed-as-miss under auto).
Either way the v2 methodology must state the policy explicitly; today's behavior is option 1 by accident, undocumented.
From the August 2026 harness audit (policy decision needed before the #139 v2 freeze).
Symptom
The response contract requires a short explanation per output, but the requirement is not enforced at scoring time: blank-explanation answers score normally.
Evidence (verified)
Decision for v2
Two coherent options:
Either way the v2 methodology must state the policy explicitly; today's behavior is option 1 by accident, undocumented.