Skip to content

v0.10.0

Latest

Choose a tag to compare

@github-actions github-actions released this 09 Oct 16:46
2339fa2

agentevals v0.10.0

agentevals now reads, stores and evaluates agent telemetry as plain OpenTelemetry. Send OTLP from any GenAI instrumentation, evaluate turns against a golden run, and get the scores back into your observability pipeline as OTel events.

Highlights

  • OpenTelemetry native ingestion. OTLP over HTTP (:4318, now spec compliant with gzip, proper 400 errors and protobuf responses) and gRPC (:4317). Files can be OTLP JSON or protobuf, Jaeger, JSONL, or Collector file exporter output. See Sending telemetry to agentevals.
  • One conversation model everywhere. The CLI, the API, MCP, the live UI and the run worker share the same turns. A model call wrapped by a framework counts once, a delegate agent's usage counts toward its turn, and tool rounds join across traces.
  • Scores as OTel events. --emit-otel (or AGENTEVALS_EVALUATION_EVENTS=true) sends every score as a gen_ai.evaluation.result log event attached to the turn it scores, so scores show up next to your traces in any backend. See OpenTelemetry pipelines.
  • Live sessions. Sessions are grouped by agentevals.session_name, gen_ai.conversation.id or session.id. Reruns get their own session, retried spans are deduplicated, and a memory budget caps the store (AGENTEVALS_LIVE_MAX_BYTES). See Sessions.
  • Live UI:
    • Failed turns are shown and counted.
    • Repeated calls to the same tool each show their own result.
    • Sessions are sorted by start time.
    • Sessions without captured message content are marked and can't become a golden run.
    • The empty page lists the OTLP endpoints.
  • Grouping control. --group-by auto|trace|conversation decides what one evaluation covers. See Grouping traces into conversations.
  • Custom evaluator protocol 1.1. Tool responses gain a structured response field; 1.0 evaluators keep working. Upgrade the SDK to use it: pip install -U agentevals-evaluator-sdk (0.2.0). See Protocol versioning.
  • Security:
    • Remote evaluator refs must be plain relative paths, and servers only run refs listed in their source's index.
    • The evaluator cache is stricter.
    • The production Collector recipe keeps message content out of your backend.
    • SDK exporters never reuse your application's TLS client certificates.

Experimental: kagent 1.0 harness agents via Substrate

  • Claude Code harness: tested on kagent v1.0.0-alpha9. With the Collector recipe in the Kubernetes example, you get tool calls by name and model calls with token usage. Tool arguments aren't recorded yet, so trajectory scores against a golden run with arguments will fail.
  • Codex harness: not tested yet. Full support depends on kagent writing runtime GenAI spans (kagent#3121).
  • Go ADK agents: full model and tool calls with arguments, results and tokens.
  • The example includes a Collector filter for kagent's TaskStore spans, which keeps long turns under the per trace span limit, plus commands to send prompts from the command line.

❗Breaking changes

Ingestion

  • WebSocket ingestion (/ws/traces) and AgentEvalsStreamingProcessor are removed. Send OTLP to :4318 or :4317. AgentEvals SDK sessions switch over on their own.
  • Receivers reject spans and logs with invalid or all zero ids and count them in partialSuccess.

Results

  • Turn boundaries and token counts follow the new model, so scores on the same telemetry can differ from v0.9.x.
  • An unmatched eval case stays NOT_EVALUATED instead of falling back to the first case. The exception: an eval set with one case evaluated against one trace group is still paired. See How traces are matched to cases.
  • invocation_id is now the turn's anchor span id.
  • Tool response output given to custom evaluators is JSON text, no longer a Python repr. Read the structured response field instead (SDK 0.2.0).

API

  • Keys inside args, arguments, response, result, output and details are no longer converted to camelCase.
  • get-trace returns one OTLP JSON document in trace_content, no longer JSONL spans.
  • The SSE span_received event carries span stubs (ids and name) only.

Evaluators

  • The openai_eval evaluator and the [openai] extra are removed, because OpenAI shuts down its Evals API on 2026-10-31. Use response_match_score, final_response_match_v2 or a code evaluator instead. Migration 000002_drop_openai_eval deletes old openai_eval result rows.

Helm

  • rbac.create: true now requires rbac.secretNames, or an explicit rbac.allowAllSecrets: true. See the chart values.

Old files

  • Files and bug report bundles written by older versions that copied log content into every span are read as concatenated user text. Record a fresh session.

Upgrading

  1. Point your exporters or Collector at :4318 (HTTP) or :4317 (gRPC). See Forward agent telemetry to agentevals.
  2. Turn on message content capture in your instrumentation. See Producer setup for per framework settings.
  3. Helm users: set rbac.secretNames or rbac.allowAllSecrets.
  4. Replace any openai_eval evaluators before upgrading the database.

What's Changed

New Contributors

Full Changelog: v0.9.12...v0.10.0