RFC: first-class hard constraints (Result.constraint_violations) — distinct from graders and errors
#47
Replies: 2 comments
|
Reference implementation is up: #48 (draft, marked DO NOT MERGE — same status as PR #46 for Discussion #45, staying open only so there's something concrete to test against during the comment period). It implements exactly what's proposed above: Still very open to pushback on field names, whether |
Status update: comment period still open (~3 more days), and this is exactly the case where the calendar shouldn't drive the decisionUnlike #45, this one isn't close to done just because two weeks are almost up (through 2026-09-23). @MSKazemi — whose own framing this entire RFC is built from — is still actively, substantively engaged as of a few days ago: a second detailed review caught a genuine correction ( So: this RFC is not being closed out or merged yet, comment period or no comment period. It'll wait for that fixture — a data point generated from a real system beats anything I'd hand-author, and rushing to merge while the person whose insight this is built on is still mid-contribution would be the wrong call regardless of what the calendar says. PR #48 stays a draft. Thanks again for the depth here, @MSKazemi — genuinely one of the best reviews this project has gotten, and worth waiting for. |
Uh oh!
There was an error while loading. Please reload this page.
Summary
This RFC exists because of a real conversation, not a self-critique item, so I want to credit it precisely before getting into the schema: while checking whether an
aobench-openeval-adaptermade sense (MSKazemi/aobench#51), AOBench's maintainer (@MSKazemi) identified a genuine gap in how this spec models evaluation outcomes — not a naming quibble, a missing concept. Everything below is my attempt to turn that into a concrete proposal; the underlying insight is theirs, and I've tried to quote it rather than paraphrase it away.Restating the gap against the current schema
resultset.jsontoday (verified by reading the live schema, not from memory) gives aResultexactly two ways to be non-passing:grader_results[]— scored, averaged-in-spirit quality signals.GraderResult.scoreis[0,1]ornull;passedis a per-grader threshold check. Nothing about this array is designed to dominate; a consumer is expected to look at the whole set.error(type: "timeout" | "provider_error" | "runner_error") — a closed three-value enum, all describing execution failures: the test case did not produce a usable result at all. Note the enum is closed and specifically about the harness/provider, not about the content of a produced answer.AOBench's real
dimension_scores(seven independently-scored axes:outcome,tool_use,grounding,governance,robustness,workflow,efficiency) map cleanly onto sevenGraderResults. But AOBench also hashard_fail: bool— an RBAC/governance violation that zeroes the entire task score regardless of how well the other six dimensions scored, because (as @MSKazemi put it) "a correct answer obtained by overstepping your authority is not a good outcome that needs a caveat. It is not an outcome at all." Neither existing field models that correctly:score: 0on agovernanceGraderResultloses the propagation. A consumer averaging seven scores where six are near-perfect and one is zero computes ≈0.85 and reports a success. @MSKazemi called this out directly as "the dangerous failure mode, because it produces a plausible number."Result.errorloses the fact that the run succeeded — it executed, produced a real answer, and was fully, legitimately scored on six real dimensions.errorexisting at all means consumers reasonably drop those rows from denominators (the schema's own framing: "the test case errored"), so a model that hard-fails constantly would get a better apparent average by having its worst rows silently excluded — inverting the metric in the opposite direction from the averaging problem above.Both are real failure modes, not hypothetical ones — a consumer using only what's in the schema today has no way to represent this case correctly.
Proposal
@MSKazemi's own suggested shape, which I've kept close to verbatim because it's already right:
Concretely, as a
spec/schemas/resultset.jsonaddition toresults[].properties:Worked example, using AOBench's real case
{ "test_case_id": "tc_rbac_001", "actual_output": "The billing overrun on cluster-07 was caused by the retry storm logged at 03:14 UTC.", "passed": false, "grader_results": [ { "grader_id": "grounding", "type": "custom", "score": 0.97, "passed": true }, { "grader_id": "outcome", "type": "custom", "score": 1.0, "passed": true } ], "constraint_violations": [ { "id": "rbac_scope", "type": "authorization", "detail": "Answer cites cluster-07 billing logs; the requesting role's scope grants read access to cluster-03 only.", "invalidates_result": true } ] }Every
grader_resultsentry is a genuine pass — the agent's actual quality was real and stays visible for debugging/analysis.constraint_violationsoverrides the outcome:passedisfalse, and this row still counts insummary.totalthe same as any other failed result, unlikeerror.Two validation rules, one of them cross-field (hand-rolled-only, same class as
SELF_PARENT/DUPLICATE_ATTEMPT)REQUIRED— each entry needs a non-emptyidand a booleaninvalidates_result. Structurally expressible in plain JSON Schema.CONSTRAINT_INVALIDATES_PASS— if any entry hasinvalidates_result: true, theResult's ownpassedMUST befalse. Likegroup.parent_group_id'sSELF_PARENTrule (Discussion RFC: grouped/sibling ResultSets with rollup semantics (mutation testing, seed sweeps, model comparisons) — follow-up to #36 #45) and(test_case_id, run_id, attempt)'sDUPLICATE_ATTEMPTrule (Discussion RFC: repetition/attempt tracking in ResultSet (raised by AgentVerity via issue #20) #22), this is a cross-field constraint the raw JSON Schema'sproperties/requiredvocabulary can't express without a$datareference, which this project's schemas deliberately avoid — so it's enforced by both reference SDKs' hand-rolled validators only, with a conformance fixture and a schema-consistency test documenting exactly that gap (the JSON Schema alone accepts the structurally-fine-but-semantically-wrong document; the hand-rolled validator correctly rejects it).A second, independent precedent — not just one benchmark's design choice
One real system is a data point, not a pattern, so I looked for a second, unrelated one before treating this as settled. Guardrails AI (
guardrails-ai/guardrails) already has a live, shipped taxonomy for exactly this distinction, confirmed against its own current docs: a validator'son_failaction is one ofreask,fix,filter,refrain,noop,exception,fix_reask,custom.fixis a soft, corrective action (substitute a corrected value and continue) — the rough analogue of a scored grader.refrain(returnNone, discard the output) andexception(raise, halt execution) are categorical — a validation failure that invalidates the whole generation regardless of what else about it was fine, independent ofnoop, which records the failure in logs but takes no action (closer to today'sGraderResult— recorded, but not overriding). That's an industry precedent, arrived at independently of AOBench and of this spec, for "some failures are categorical/dominant, not partial-score signals" — applied at generation-guardrail time rather than post-hoc eval scoring, but the same underlying distinction @MSKazemi is pointing at for evaluation results specifically.What this deliberately does not do
It does not define a closed vocabulary for constraint
types orids — those are as domain-specific asgroup.role(Discussion #45) orisolation(Discussion #22), so both stay open strings. It also does not change howsummaryis computed —summaryis already producer-computed and optional; a producer that wantssummary.avg_scoreto reflectconstraint_violationsis free to compute it that way, but this schema doesn't mandate a formula, the same restraint already applied to not standardizing agrouprollup.Why not the two options actually on the table before this?
GraderResult.score = 0— rejected above: loses the "this run genuinely succeeded on six other dimensions" signal and, worse, launders into a plausible-looking averaged score that's actually wrong.Result.error— rejected above:error's own three-value enum is specifically about execution/harness failures, and its semantics ("no usable result") get the denominator effect backwards for a constraint violation, which produced a perfectly usable, gradable result.Process
Same as every prior spec change here (Discussions #8–#11, #22, #45): two-week comment period in this Discussion first. I've already built and locally verified a complete reference implementation (schema, both SDKs, both dual-validation-path cross-check suites, four new conformance fixtures, and a
spec/SPEC.mdExtension Mechanism subsection) on a draft PR, explicitly marked DO-NOT-MERGE for the same reason PR #46 is — so there's something concrete to review and test against, not a claim that consensus already exists. Link to follow once the PR is up. Very open to pushback on the field names, whethertypeneeds any conventional vocabulary, or whether the two-rule validation is the right shape — @MSKazemi, tagging you specifically since this is fundamentally your framing, not mine; I'd rather you correct the shape now than after it's harder to change.All reactions