Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

11 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Shanghai AI Lab

A2E   An End-to-End Agent Auditing Engine

Evaluate any agent on any dataset, with full trajectory visibility.

English | 中文


A2E (Agent Auditing Engine) is an open-source platform for auditing agent harnesses end to end. It helps you:

  • Build experiments — pair any benchmark with any agent harness and run them through one CLI
  • Capture trajectories — auto-instrument LLM and tool calls into standardized traces
  • Score results — score both the process and the final outcome with multidimensional metrics
  • View results — browse datasets, experiments, and trace trees in a local UI

The loop is simple: build → capture → score → view. A local server stores runs, traces, and scores for every step.

A2E pipeline

Contents

  1. Quick start
  2. Build experiments
  3. Capture trajectories
  4. Score results
  5. View results

1. Quick start

bash script/start.sh

Prerequisites are listed in the script header. This syncs dependencies, builds the UI, creates .env from .env.example when missing, and starts the server at http://localhost:6006.

Before running experiments, fill in API keys in .env. Field meanings and override priority are documented in .env.example.

2. Build experiments

Pick a dataset and an agent harness, then run:

# Terminal 1 — keep the server up
bash script/start.sh                 # → http://localhost:6006

# Terminal 2
cd task
set -a; . ../.env; set +a
export no_proxy="127.0.0.1,localhost,${no_proxy:-}"; export NO_PROXY="$no_proxy"
uv run --frozen python examples/run_experiment.py \
  --dataset tau-bench --agent agno --model qwen-max \
  --evaluators tool_recall,llm_judge --domain retail

Or follow the interactive walkthrough: bash example/run_examples.sh.

CLI

uv run --frozen python examples/run_experiment.py \
  --dataset <dataset> \
  --agent   <framework>      # default: agno
  --model   <model>          # default: A2E_MODEL from .env
  --evaluators <a,b,...>     # default: exact_match,substring
  --n       <count>          # default: 40
  --sample-seed 20260721     # optional, reproducible sample
  --domain  retail           # tau-bench / tau2 only
Flag Purpose
--list Print all datasets / agent harnesses / evaluators
--dataset Dataset name (required)
--agent Agent harness (default agno)
--model Override A2E_MODEL for this run
--evaluators Comma-separated scorers
--n Sample size (absolute count, random without replacement)
--domain retail / airline (tau-bench / tau2)
uv run --frozen python examples/run_experiment.py --list

Datasets & agent harnesses

Datasets
tau-bench, tau2, tau3, traject-bench, mmlu, gsm8k, humaneval, gpqa, mmlu-pro, math, bbh, swe-bench-lite, swe-bench-verified, swe-bench-pro, terminal-bench-2, terminal-bench-2.1, …
Agent harnesses
smolagents · agno · llama-index · langgraph · crewai · google-adk · autogen-agentchat · claude-sdk · openai-agents

Built-in scorers include exact_match, substring, tool_recall, numeric_match, mc_letter, swe_resolved, swe_fail_to_pass, swe_pass_to_pass, tb_resolved, and llm_judge. Recommended combos live in task/runners/.../registry.py (default_evaluators).

Sandbox experiments

Sandbox datasets need Docker and pullable images (1–3 GB each). Pin a cached instance before formal runs:

A2E_SWE_INSTANCE="<cached_instance_id>" \
uv run --frozen python examples/run_experiment.py \
  --dataset swe-bench-lite --agent agno --n 1 \
  --evaluators swe_resolved,swe_fail_to_pass,swe_pass_to_pass
  • swe-bench-proA2E_SWE_PRO_INSTANCE
  • terminal-bench-2A2E_TB2_TASK=<task>

3. Capture trajectories

While an experiment runs, Monitor auto-instruments the agent: LLM calls, tool calls, and related spans are written to the server as OpenTelemetry / OpenInference traces. You do not need a separate capture step for supported harnesses.

Browse captured trajectories at http://localhost:6006 under Experiment-<id> projects:

Page What you see
/datasets Uploaded samples
/experiments Per-sample scores
/projects Trace trees (LLM + tool calls)

4. Evaluation

Evaluate agent runs from both the process and outcome perspectives.

Once trajectories are collected, the evaluation module analyzes agent behavior, including planning, tool usage, memory, efficiency, safety, and final task correctness.

Evaluations can be run at different granularities:

  • Run the complete evaluation suite to obtain an overall profile.
  • Run individual evaluation groups to analyze specific capabilities.

Run all evaluations

# Start the evaluation service
bash script/start.sh

# Run the full evaluation pipeline
cd server
uv run python ../eval/scripts/run_eval.py \
  --base-url http://localhost:6006 \
  --experiment-id <id> \
  --part all
## 5. View results

The React viewer is served by `a2e serve` at http://localhost:6006 (built by `script/start.sh`). Swipe between **Task / Trace / Eval** for the same sample; open Trace for the span tree.

### Development (HMR)

`script/start.sh` builds the production UI. For Vite HMR, start the API and the UI separately:

```bash
# Terminal 1 — API only (after deps are installed)
cd server && uv run a2e serve       # http://127.0.0.1:6006

# Terminal 2
cd ui && pnpm install && pnpm dev   # http://127.0.0.1:5173  (proxies /v1 → :6006)

Or use a2e serve --dev for Vite HMR via the server templates.

Project layout

AEP/
├── script/              # One-click install + server start
├── task/                # Build experiments (datasets, agents, runners)
├── monitor/             # Capture trajectories (auto-instrumentation)
├── eval/                # Score results (process and outcomes)
├── server/              # Store runs, traces, and scores
├── ui/                  # View results
└── example/             # Interactive walkthrough

Acknowledgements

A2E depends on and draws from the following open-source projects:

About

No description, website, or topics provided.

Resources

Stars

17 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages