RFC: distinguishing FAILED from UNVERIFIED — a Result can fail without proving it failed #49
Replies: 3 comments
Reference implementation is up: #83 (draft, DO NOT MERGE)This thread has had no comments since it opened on 2026-09-11, so I built the concrete version to give it something to react to. #83 is a draft reference implementation, not a proposal to merge. It has the same status as #46 (for #45), #48 (for #47) and #66 (for #58 / #67). It covers the JSON Schema (plus its mirror), both SDKs, both schema-consistency suites, 5 conformance fixtures, and a PROPOSED subsection in The comment period restarts today, 2026-09-27, and runs two weeks, through 2026-10-11. Nothing merges before then. If there is rough consensus after that, the implementation gets re-opened fresh against What #83 proposesOne optional field,
Rules, all three enforced by the JSON Schema itself (
Both SDKs check the full 28-case truth table against both validation paths. The raw schema and the hand-rolled validators agree on every case. Consumer semantics in the spec text:
How it relates to what the spec already hasValidation Rule 6 already has a per-grader "not verified":
So "all scores null ⇒ unverified" is a SHOULD, not a validator rule. Composition with the other open RFCs is documented but not implemented, because those fields are not on
Alternatives rejected
A correction to my opening postThe "Retry safety" bullet above has it backwards. I wrote that an UNVERIFIED case is safe to redo and a FAILED one is not. @soul-sol's framing is the opposite, and it is the correct one: FAILED is "provably didn't happen, safe to re-run", and UNVERIFIED "may have happened", so blindly retrying it can re-execute side effects. #83's spec text follows soul-sol's framing. What I found in agent-watchI read What the repo does implement is the launcher-records-the-exit-code half:
One fixture uses the verbatim Questions for @soul-sol
Pushback on the field name, the rules, or the whole shape is welcome, and so are counter-examples from other runtimes. |
Status, 2026-09-30: still no outside comments, and one new data pointStill no comment from anyone outside the project (the comment period runs through 2026-10-11, and nothing merges before then). Two things to put on the table so the thread has more than my own writing in it. A data point from building the Every Eval Ever exporter ( I've asked one outside group for its view. With the approval of the project's maintainer, I emailed the AAIF Accuracy & Reliability working group's list today with a short introduction and one question: does reliability and accuracy reporting need a first-class "not verified" outcome, and what should it be called. I stated that the project is small and has no production adopters. If anyone from that group or elsewhere lands here, the four questions in my 2026-09-27 comment are still the useful place to start, and so is a counter-example from a runtime where No change to the proposal or to the dates. |
Status, 2026-10-01: first outside data point, proposed narrowing, and the bar for landingNo change to the dates (comment period still runs through 2026-10-11). Three things for the thread. 1. A data point from a second runtime. Over on AgentEval#203, one of AgentEval's maintainers listed the result states its framework distinguishes: 2. Proposed narrowing of #83, for objection.
3. The bar for landing, and the alternative it has to beat. I don't want this to merge because a calendar date passed. My proposal is that #83 only goes forward if at least one implementer other than me produces a real fixture, or says in this thread that their system needs the field. If nobody does by 2026-10-11, I'll extend the period rather than merge, and be ready to close the RFC as a documented convention: Rule 6 per grader, plus a A related point for anyone designing against this: #47 and #49 both change what a non-clean Result means. Before either lands I'll write a one-page outcome-model table (passed / failed / error / unverified / constraint-violated) so the two don't drift into overlapping fields. The four questions in my 2026-09-27 comment are still the useful place to start. Objections to any of the above are welcome, including "the field isn't worth it". |
Uh oh!
There was an error while loading. Please reload this page.
Where this comes from
Raised by @soul-sol in a semantic-kernel discussion about a portable EvalPort adapter for multi-agent orchestrations, responding to a point about how badly
str(final_output)collapses multi-agent runs. Quoting the part that matters here directly, because it's already right:The concrete case: an orchestration leaves a worker agent in a terminal state of "alive, no output, no question" — not finished, not asking for input, just unresolved. That's not the same thing as an orchestration that ran to completion and produced a wrong answer, and it's not the same thing as a harness-level execution failure either.
Why the existing schema can't say this today
A
Resulthas exactly two ways to be non-passing, and neither is "I don't know":grader_results[]+passed: false— a grader looked at the output and judged it wrong. This asserts the run was judged and failed the judgment.error(timeout | provider_error | runner_error) — the harness itself failed to produce a usable result. This asserts nothing ran to completion."The run executed, produced some terminal state, and no grader can honestly say whether it succeeded" is neither of those. Forcing it into
passed: falseclaims a judgment that was never made. Forcing it intoerrorclaims an execution failure that didn't happen — the orchestration didn't crash, it just ended somewhere ambiguous. Today's producers are stuck picking the less-wrong lie.This is the same class of gap as Discussion #47 (
Result.constraint_violations, from a different maintainer's benchmark, currently in its own comment period with a reference PR at #48): a real distinction that the existing two-axis model (grader judgment vs. harness error) structurally cannot express, surfaced by someone who actually ran evals against real fixtures and hit the wall rather than by guessing at the schema in the abstract.Two concrete consequences of collapsing them
Both from soul-sol's own writeup, and both are consumer-facing problems, not just modeling aesthetics:
What this is not (yet) a proposal for
I don't have a schema shape I'm confident in — this Discussion is the "does this gap exist and is it worth solving in EvalPort specifically" step, per this project's own governance process (two weeks of comment before anything gets implemented, the same path #45 and #47 went through). Two rough directions, neither committed to:
Resultdistinguishing at minimumfailed/unverified/ (implicitly)passed, sitting alongside the existingpassed: boolrather than replacing it —passedstays the grader-judgment axis, this would be a new, separate "was a judgment even possible" axis.erroras a fourth type (e.g.error.type: "unverifiable") — cheaper structurally, but risks the same denominator-exclusion problem RFC: first-class hard constraints (`Result.constraint_violations`) — distinct from graders and errors #47 already identified forerror: consumers reasonably droperrorrows from denominators, and an UNVERIFIED result is emphatically not something that should silently vanish from a report the way a harness crash should.soul-sol's own three-state terminal taxonomy for multi-agent workers specifically (
finished/needs-input/unknown, recorded by the launcher rather than scraped from the transcript) is a good concrete grounding case if this goes anywhere — it's a real, independently-shipped implementation (MIT, soul-sol/agent-watch), not a hypothetical.Also worth surfacing separately from the schema question, since it's good practice regardless of what happens here: soul-sol's writeup names three ways an eval suite can pass while being wrong (positive-only fixtures with no case that must fail, a suite that only proves the mapping executes rather than that the mapping is right, a threshold that silently discards valid data). None of that needs a spec change — it's a testing-discipline note for anyone (including me) writing conformance fixtures for an EvalPort adapter, and I'm folding it into how I build the next one.
@soul-sol — tagging you since this is your framing and your worked example; corrections welcome before this goes further, especially if either rough direction above already looks wrong from where you're sitting.
All reactions