Skip to content

Explanation contract is unenforced — set the v2 scoring policy #145

Description

@MaxGhenis

From the August 2026 harness audit (policy decision needed before the #139 v2 freeze).

Symptom

The response contract requires a short explanation per output, but the requirement is not enforced at scoring time: blank-explanation answers score normally.

Evidence (verified)

  • 61 blank-explanation rows accepted on the current board: grok-4.3 = 56, kimi-k2.6 = 4, haiku-4.5 = 1.
  • If blanks scored as misses, grok-4.3 would drop 2.574 pts (~7 ranks). No other model moves materially.

Decision for v2

Two coherent options:

  1. Score on numbers only (recommended): keep the explanation request in the prompt, publish explanation-compliance as a per-model column, but don't gate the accuracy score on prose compliance. Scoring prose mixes constructs into a numeric-accuracy headline.
  2. Enforce as miss, stated up front alongside the parse-failure policy (like Sonnet 5's 56/1,984 unparsed-as-miss under auto).

Either way the v2 methodology must state the policy explicitly; today's behavior is option 1 by accident, undocumented.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions