Repository navigation
agentevals v0.10.0
agentevals now reads, stores and evaluates agent telemetry as plain OpenTelemetry. Send OTLP from any GenAI instrumentation, evaluate turns against a golden run, and get the scores back into your observability pipeline as OTel events.
Highlights
- OpenTelemetry native ingestion. OTLP over HTTP (:4318, now spec compliant with gzip, proper 400 errors and protobuf responses) and gRPC (:4317). Files can be OTLP JSON or protobuf, Jaeger, JSONL, or Collector file exporter output. See Sending telemetry to agentevals.
- One conversation model everywhere. The CLI, the API, MCP, the live UI and the run worker share the same turns. A model call wrapped by a framework counts once, a delegate agent's usage counts toward its turn, and tool rounds join across traces.
- Scores as OTel events.
--emit-otel(orAGENTEVALS_EVALUATION_EVENTS=true) sends every score as agen_ai.evaluation.resultlog event attached to the turn it scores, so scores show up next to your traces in any backend. See OpenTelemetry pipelines. - Live sessions. Sessions are grouped by
agentevals.session_name,gen_ai.conversation.idorsession.id. Reruns get their own session, retried spans are deduplicated, and a memory budget caps the store (AGENTEVALS_LIVE_MAX_BYTES). See Sessions. - Live UI:
- Failed turns are shown and counted.
- Repeated calls to the same tool each show their own result.
- Sessions are sorted by start time.
- Sessions without captured message content are marked and can't become a golden run.
- The empty page lists the OTLP endpoints.
- Grouping control.
--group-by auto|trace|conversationdecides what one evaluation covers. See Grouping traces into conversations. - Custom evaluator protocol 1.1. Tool responses gain a structured response field; 1.0 evaluators keep working. Upgrade the SDK to use it:
pip install -U agentevals-evaluator-sdk(0.2.0). See Protocol versioning. - Security:
- Remote evaluator refs must be plain relative paths, and servers only run refs listed in their source's index.
- The evaluator cache is stricter.
- The production Collector recipe keeps message content out of your backend.
- SDK exporters never reuse your application's TLS client certificates.
Experimental: kagent 1.0 harness agents via Substrate
- Claude Code harness: tested on kagent v1.0.0-alpha9. With the Collector recipe in the Kubernetes example, you get tool calls by name and model calls with token usage. Tool arguments aren't recorded yet, so trajectory scores against a golden run with arguments will fail.
- Codex harness: not tested yet. Full support depends on kagent writing runtime GenAI spans (kagent#3121).
- Go ADK agents: full model and tool calls with arguments, results and tokens.
- The example includes a Collector filter for kagent's
TaskStorespans, which keeps long turns under the per trace span limit, plus commands to send prompts from the command line.
❗Breaking changes
Ingestion
- WebSocket ingestion (
/ws/traces) and AgentEvalsStreamingProcessor are removed. Send OTLP to :4318 or :4317. AgentEvals SDK sessions switch over on their own. - Receivers reject spans and logs with invalid or all zero ids and count them in partialSuccess.
Results
- Turn boundaries and token counts follow the new model, so scores on the same telemetry can differ from v0.9.x.
- An unmatched eval case stays
NOT_EVALUATEDinstead of falling back to the first case. The exception: an eval set with one case evaluated against one trace group is still paired. See How traces are matched to cases. - invocation_id is now the turn's anchor span id.
- Tool response output given to custom evaluators is JSON text, no longer a Python repr. Read the structured
responsefield instead (SDK 0.2.0).
API
- Keys inside args, arguments, response, result, output and details are no longer converted to camelCase.
get-tracereturns one OTLP JSON document intrace_content, no longer JSONL spans.- The SSE span_received event carries span stubs (ids and name) only.
Evaluators
- The openai_eval evaluator and the
[openai]extra are removed, because OpenAI shuts down its Evals API on 2026-10-31. Use response_match_score, final_response_match_v2 or a code evaluator instead. Migration000002_drop_openai_evaldeletes oldopenai_evalresult rows.
Helm
rbac.create: truenow requiresrbac.secretNames, or an explicitrbac.allowAllSecrets: true. See the chart values.
Old files
- Files and bug report bundles written by older versions that copied log content into every span are read as concatenated user text. Record a fresh session.
Upgrading
- Point your exporters or Collector at :4318 (HTTP) or :4317 (gRPC). See Forward agent telemetry to agentevals.
- Turn on message content capture in your instrumentation. See Producer setup for per framework settings.
- Helm users: set rbac.secretNames or rbac.allowAllSecrets.
- Replace any openai_eval evaluators before upgrading the database.
What's Changed
- chore(deps): bump the python group with 4 updates by @dependabot[bot] in #235
- fix: harden remote evaluators by @krisztianfekete in #237
- refactor: unship the OpenAI Evals and fix baseline migration rollback by @krisztianfekete in #238
- fix(otlp): make the OTLP/HTTP receiver spec compliant by @JHf0912 in #223
- feat: OTel-native refactor by @krisztianfekete in #241
- test: kagent compat fixtures, ordered replay, and a daily Weaver check of kagent's registry by @krisztianfekete in #245
- fix: keep GenAI content out of the backend recipe, isolate SDK TLS, set content capture by @krisztianfekete in #246
- fix: smaller release blockers by @krisztianfekete in #247
New Contributors
Full Changelog: v0.9.12...v0.10.0