Repository navigation
RFC: rate-based GraderResult denominators (trials/successes) — promoted from #58 #67
adhabnr-ux
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
This promotes issue #58 to a full RFC and opens the two-week comment period that this project's Governance section requires before any schema change. The question is narrow: should a rate-based
GraderResult.scorecarry its denominator, and if it should, where does it live and is it required?A reference implementation of one option (Option A below) is on a draft, do-not-merge PR so there is something concrete and tested to review: #66. It is not a proposal to merge ahead of this discussion.
Where this came from
@sattyamjjain raised this while we were discussing whether a
garak-openeval-adaptermade sense (NVIDIA/garak#2110). He maintains a red-team harness for robot vision-language-action policies and is not affiliated with garak or EvalPort. The framing is his:The discussion started in #58, which asked whether the gap was real enough to warrant a schema change at all. This Discussion is the next step in the same process that #22 and #45 went through.
The gap
GraderResult.scoreinspec/schemas/resultset.jsonis a bare number in[0, 1], ornull. Some graders are reallysuccesses / trials: pass rate over N trials, attack success rate over N attempts. For those graders the schema has no place for N. A score of 0.0 from 0 successes in 5 trials and a score of 0.0 from 0 in 500 serialize identically, even though their statistical support differs by two orders of magnitude.Today a producer that wants to record N has to put it in
GraderResult.metadataunder a key of its own choosing. No generic consumer knows to look for it, and a producer that forgets produces a document that looks complete and is not.Why
Result.attemptdoes not already cover thisWhen each trial is its own
Resultrow, N is already derivable: count the rows that share(test_case_id, run_id)usingResult.attempt(Discussion #22). That is the right model for harnesses that emit one row per trial.It does not cover a single
GraderResultthat aggregated N trials internally. For example, a grader might run 500 episodes of one attack and report one attack-success rate. There is oneResultand oneGraderResult, so there are no rows to count. The denominator exists only inside the grader, and the schema currently has no place for it. This RFC is about that case.Options
Option A: optional
trials/successesonGraderResult(implemented in #66)trials: integer >= 1, the denominator. Optional.successes: integer >= 0, the numerator. Optional. If present, it requirestrials, expressed in the JSON Schema asdependentRequired: {"successes": ["trials"]}, so the raw schema and the hand-rolled SDK validators agree (REQUIRED).successes <= trials(SUCCESSES_EXCEED_TRIALS) and, whenscoreis non-null,|score - successes/trials| <= 1e-6(RATE_SCORE_MISMATCH). These are cross-field rules that plain JSON Schema cannot express without$data. Both reference SDKs enforce them in their hand-rolled validators only, the same class of rule asSELF_PARENTandDUPLICATE_ATTEMPT. Each SDK's schema-consistency tests document the gap explicitly.Pros: purely additive (MINOR). Every existing document stays valid. The field names and semantics are the same across both SDKs. Consumers get a standard place to look for N.
Cons: optional fields do not fix the failure mode #58 is most concerned with. In @sattyamjjain's words: "If a
ResultSetentry can serialise a score without its denominator, every adapter author has to remember to carry it, and the ones who forget produce something that looks complete and is not." Option A makes carrying N possible and standard, but not mandatory.Option B: Option A, plus
trialsrequired when the grader opts in as rate-based (not implemented)A
Graderdefinition declares itself rate-based, for example with a grader-level flag or a reserved key such asopeneval.rate_based: true. AGraderResultfor that grader withouttrialsis then rejected.Pros: provides the enforcement #58 asks for, scoped to the graders where a denominator means something. Single-shot graders (
exact_match,json_schema, a singlellm_judgecall) are unaffected.Cons: the check becomes a cross-document rule.
validate_result_set()validates aResultSetalone and does not receive the suite'sGraderdefinitions, so either its signature changes or the "rate-based" declaration has to be copied onto eachGraderResult, where it could drift. It also means deciding where the declaration lives: a new top-level field onGrader(whose schema isadditionalProperties: false),params, or a reservedmetadatakey. It still depends on the producer setting the flag. A producer that forgets the flag is in the same position as one that forgetstrialsunder Option A, though the omission is at least made once per grader rather than once per result.Option C: no schema change, a
metadata.openeval.trialsconventionDocument a reserved
GraderResult.metadata.openeval.trials(andsuccesses) key convention inSPEC.md, with no change to the schema.Pros: no schema change is needed at all.
GraderResult.metadatais alreadyadditionalProperties: true, whereasGraderResultitself isadditionalProperties: false, which is why Option A has to change the schema. It is the smallest possible change and costs nothing to adopt or ignore.Cons: it is optional by convention, which is exactly the thing that #58 identifies as failing silently. A validator cannot type-check a convention inside open
metadatawithout special-casing it, so it would not catchtrials: "500",successes > trials, or a score that contradicts its own counts. A misspelled key (metadata.openeval.trails) validates cleanly.Explicit non-goals
successes, trials -> intervalutility with no framework dependency. That utility and this schema question are independent and do not share a review timeline.trialsalone supports a valid interval. Per @sattyamjjain's caveat, for correlated or clustered trials (e.g. episodes within one task) a naive interval over pooled trials is too narrow.trialslets a reader tell a well-supported rate from a noisy one at a glance. Nothing in the spec text should imply more than that.summaryis computed.summaryremains producer-computed and optional.Comment period
Opens 2026-09-27, closes 2026-10-11. The reference-implementation PR stays a draft throughout. If the discussion reaches rough consensus, the implementation will be re-opened fresh against whatever
mainlooks like at that point, rather than merged as-is.Questions for commenters
ResultSetvalidation) acceptable, or should the rate-based declaration be carried perGraderResult?Grader, aparamsconvention, or a reservedmetadatakey?trials-only (nosuccesses) a useful state to allow, given thatscore * trialsalready implies the numerator? Or should the two be required together so that the numerator is never reconstructed from a rounded float?RATE_SCORE_MISMATCH? Should the rule be dropped entirely in favor of treatingscoreas authoritative?Resultper trial withattempt, should the spec say anything about the relationship between a per-rowattemptcount and atrialsvalue on an aggregating grader, or leave the two independent?All reactions