Idea: aobench-openeval-adapter for portable eval interchange (EvalPort) #51
Replies: 4 comments 4 replies
|
This is a careful piece of work — you read the actual schemas rather than the README, the Taking the open question first, because it is the interesting one and my answer is An RBAC hard fail is not a grader result, and it is not an error eitherYou framed it as a choice between a scored grader with Concretely, in AOBench a governance hard fail does not set governance to zero. It zeroes So the two options you listed lose different halves of that:
What I would suggest instead: a hard constraint is a first-class, typed property of constraint_violations: list[ConstraintViolation] = []
# ConstraintViolation: {id, type, detail, invalidates_result: bool}with the rule that The general shape this points at, beyond AOBench: evals increasingly have properties On the mapping itselfMostly right. Four corrections, all small, and one of them is a bug you caught by
Which repo it should live inYour call, and I would gently suggest your repo, not ours — at least first. An exporter What I am glad to do: keep the schemas stable and documented, review the mapping properly, Two things that will make it much easier for you:
No obligation to build any of it — but if you do, I will review it properly and link it. One optional aside: if AOBench turns out to be useful as a stress case for the schema, a |
|
A short follow-up, and it is about credit rather than the design question. I went back through every account that has ever touched this repository — issues, pull You are now on the contributor I have also added a policy line saying explicitly that this kind of contribution counts. Two notes, both practical:
The offer to review an adapter against either repo stands, and there is no deadline on it. |
|
This is thorough — all five landed cleanly, and the On the fixture: still yes, and I'd rather send you something real and independently verified than something quick. I'll generate an actual RBAC-hard-fail Thank you for building this in the open the way you have — the draft-marked-DO-NOT-MERGE, two-week comment period is a real governance process, not a formality, and it shows. |
|
Thank you — and right back at it. This has been one of the most technically substantive discussions this project has had, and the constraint-modeling framing (a hard fail as a property of the row's validity, not a graded dimension) is one I expect to keep citing well beyond this adapter. I'll follow up on this thread once the fixture is ready, generated from a genuine RBAC violation and verified against AOBench's own governance scorer rather than either of the known own-job corpus defects. No date attached — I'd rather get it right than get it soon. |
Uh oh!
There was an error while loading. Please reload this page.
Hi — not a maintainer or contributor here, just looking at benchmarks with real trace-based scoring. I'm working on EvalPort, an open interchange schema for portable LLM eval data (
TestCase/Grader/Result/ResultSet/GraderResult, Python + TS SDKs).AOBench's
TaskSpec/Trace/BenchmarkResult(src/aobench/schemas/task.py,trace.py,result.py) map onto EvalPort'sTestCase/Resultclosely enough that I think a converter is worth raising — not because AOBench needs EvalPort, but because AOBench is the first RBAC-hard-fail, HPC-domain benchmark I've looked at, and it stress-tests something the schema hasn't had to answer yet: what happens when a governance violation should zero out an otherwise-correct answer.Mapping
TaskSpec→TestCasetask_id→idquery_text→inputeval_criteria.gold_answer→expected_outputgold_evidence_refs+eval_criteria.required_evidence_refs→retrieval_contexteval_criteria.expected_tool_sequence[].tool_name→expected_toolsrole,qcat,difficulty,environment_id,access_tier→metadataTrace+BenchmarkResult→Resulttask_id→test_case_idtrace.final_answer→actual_outputlatency_seconds * 1000→duration_mshard_fail→passed=False,error={"type": "rbac_violation", "reason": hard_fail_reason}— not a scored graderdimension_scores.{outcome,tool_use,grounding,governance,robustness,workflow,efficiency}→ oneGraderResultper dimension,grader_id=<dimension>,type="custom",score=valuerun_id,weight_profile_name,adapter_name,model_name,cost_estimate_usd→Result.metadata/ResultSet.metadataSketch (shape mirrors your own
BaseExporter, and the existinglangfuse-openeval-adapter— the one adapter in the EvalPort repo I checked has real code behind it, not a stub):That's the same shape whether it lands as a new package under
evalport/adapters/aobench-openeval-adapter, or as anEvalPortExporter(BaseExporter)next toLangfuseExporterinsrc/aobench/exporters/— the latter is probably lower-friction for you since it reuses the exporter registration you already have. Happy to write it as a PR against either repo rather than asking you to.One open question I don't have a good answer for:
GraderResult.scoreis required in[0, 1], which your seven dimensions already satisfy — but there's no dedicated slot for "this dimension hard-failed the whole task" other than theResult-levelerrorI used above. If you have views on whether an RBAC-style hard constraint should be a scored grader (score=0) vs. an out-of-band error, I'd like to hear them — it'd shape how EvalPort models hard constraints generally, not just for this adapter.No pressure to build this — flagging it because the shapes already line up unusually well.
— Sahi, independent contributor (not affiliated with this project)
All reactions