Skip to content

v0.2.0 — Enterprise Agent Eval Lab

Latest

Choose a tag to compare

@SayedAbbas SayedAbbas released this 07 Sep 15:02

Enterprise Agent Eval Lab v0.2.0

A versioned starting point for the Phase 2 vendor-neutral evaluation harness.

Included

  • Adapters for OpenAI, Anthropic, Gemini, and Grok, plus a deterministic local mock.
  • Seven scoring dimensions: task success, groundedness, tool selection, tool arguments, safety, escalation, and latency.
  • JSON reports, readable Markdown summaries, and a cross-provider comparison command.
  • README architecture diagram, complete clone-and-install instructions, and a reproducible mock demonstration with sample report links.

Get started

Requires Python 3.10 or later.

git clone --branch v0.2.0 https://github.com/SayedAbbas/enterprise-agent-eval-lab.git
cd enterprise-agent-eval-lab
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
agent-eval run --dataset datasets/claims.jsonl --provider mock --output results/demo.json

On Windows PowerShell, activate with .venv\Scripts\Activate.ps1 instead.

Validation and limits

  • Ruff: passed.
  • Existing test suite: 4 passed.
  • Local mock demo: 3 synthetic claims cases, overall score 1.000.

The mock constructs responses from expected answers. Its score demonstrates the evaluation/reporting pipeline, not live-model quality. Live provider calls were not run for this release. Provider availability and model identifiers must be verified for your account. This release does not establish production readiness or a universal model ranking.