Proposal: EvalPort adapter for ifixai.api.run_inspections() / TestRunResult #104
Replies: 2 comments
|
Hey Sahi, thank you for reading our codebase so closely. The mapping you sketched seems right. |
|
Thanks, Nikos — good catch on the statuses. I'll map I'll host it under EvalPort's |
Uh oh!
There was an error while loading. Please reload this page.
Hi — I maintain EvalPort, an open interchange spec + SDK (
evalport-sdkon PyPI,openeval.validate.validate_result_set()) for portable LLM eval results, with ~35 framework adapters so far. I readifixai/core/types.py,ifixai/harness/base.py, anddocs/python-api.mdand think iFixAi's result model is one of the cleanest fits I've seen for this.Why it fits:
ifixai.api.run_inspections(...)(async) returns aTestRunResult(pydantic) withtest_results: list[TestResult],overall_score,grade,passed,category_scores. EachTestResultcarriestest_id,score,threshold,passed,status(TestStatus), andevidence: list[EvidenceItem], whereEvidenceItemhasprompt_sent,expected/expected_behavior,actual/actual_response,passed, and (for judge-scored items)judge_verdict/dimension_scores/rubric_weighted_score. That's already a per-test-case input/expected/actual/pass-fail record — exactly what OpenEval's ResultSet wants, at theEvidenceItemlevel, rolled up per inspection at theTestResultlevel.Sketch (before → after):
The adapter would map each
EvidenceItemto an OpenEval result row (test_case_id → id, prompt_sent → input, expected_behavior → expected, actual_response → actual, passed/rubric_weighted_score → score+verdict), and eachTestResult'stest_id/category/threshold/status(includingINCONCLUSIVE/ERROR, not just pass/fail) into the ResultSet's per-suite grouping, withTestRunResult.overall_score/gradeas the summary.I'd like to build
adapters/ifixai-openeval-adapter/(standard layout —pyproject.tomlwith a self-referencing pinned extra per the EvalPort convention,src/ifixai_openeval_adapter/__init__.pyexposingto_openeval()/results_to_openeval(),tests/test_adapter.py,README.md) and submit it as a PR, either here or under EvalPort'sadapters/directory — whichever you'd rather host it in.One question before I start: since
statusincludesINCONCLUSIVEandERRORalongsidePASS/FAIL, would you want those mapped to askipped/errorOpenEval verdict, or excluded from the portable set entirely (matching howinsufficient_evidencealready excludes them from your own scoring denominator)? Happy to follow whichever convention you'd prefer.— Sahi (independent contributor, not affiliated with iFixAi)
All reactions