Enterprise Agent Eval Lab v0.2.0
A versioned starting point for the Phase 2 vendor-neutral evaluation harness.
Included
- Adapters for OpenAI, Anthropic, Gemini, and Grok, plus a deterministic local mock.
- Seven scoring dimensions: task success, groundedness, tool selection, tool arguments, safety, escalation, and latency.
- JSON reports, readable Markdown summaries, and a cross-provider comparison command.
- README architecture diagram, complete clone-and-install instructions, and a reproducible mock demonstration with sample report links.
Get started
Requires Python 3.10 or later.
git clone --branch v0.2.0 https://github.com/SayedAbbas/enterprise-agent-eval-lab.git
cd enterprise-agent-eval-lab
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
agent-eval run --dataset datasets/claims.jsonl --provider mock --output results/demo.jsonOn Windows PowerShell, activate with .venv\Scripts\Activate.ps1 instead.
Validation and limits
- Ruff: passed.
- Existing test suite: 4 passed.
- Local mock demo: 3 synthetic claims cases, overall score 1.000.
The mock constructs responses from expected answers. Its score demonstrates the evaluation/reporting pipeline, not live-model quality. Live provider calls were not run for this release. Provider availability and model identifiers must be verified for your account. This release does not establish production readiness or a universal model ranking.