Skip to content

v0.10.1-beta — Samples Consolidation + Generic Renderers + Canonical Store

Choose a tag to compare

@joslat joslat released this 18 May 15:12
· 216 commits to main since this release
c8e7e7d

Adds a uniform IEvalResultRenderer contract (HTML + PDF) in AgentEval.Abstractions, with implementations in AgentEval.Core (HTML) and the new AgentEval.Rendering.Pdf project. Consolidates the per-family *.Demo projects into one focused samples/AgentEval.Samples/Benchmarks/ suite — 10 examples wired as menu group H. Every running sample writes the canonical run through FileSystemOutputStore to the repo-root .agenteval/ workspace so Mission Control + agenteval doctor see it, plus a sidecar output/{family}/run-{utc}-{suffix}/ for direct human consumption of HTML + PDF.

End-to-end verified against a real Azure agent on gpt-5-chat before merge; full audit-chain integrity (manifest contentHash ↔ HTML/PDF footer ↔ compliance evidence sourceRun.manifestHash) confirmed.

Highlights

  • IEvalResultRenderer uniform rendering contract in AgentEval.Abstractions — RenderAsync(EvalResult, options) → byte[]. Framing metadata (subject, run id, audit hash, AgentEval version) flows through EvalResultRenderOptions.
  • HtmlEvalResultRenderer (Core) — self-contained HTML, inline CSS, <details> collapsibles, XSS-safe via WebUtility.HtmlEncode, severity- and label-coded badges (threshold-fail composites no longer render green), honest NOT TESTED for skipped leaves on both leaves AND the cover banner.
  • AgentEval.Rendering.Pdf new project with PdfEvalResultRenderer — QuestPDF-backed; cover page + component summary + per-leaf detail pages + audit-chain appendix. Embedded into the umbrella NuGet via PrivateAssets="all".
  • samples/AgentEval.Samples/Benchmarks/ sample suite — 10 focused examples in menu group H: Registry Discovery, Performance, Agentic, GDPR, EU AI Act, OWASP, MITRE, LongMemEval, Memory, Report Browser.
  • Canonical store wiring — every running sample writes a canonical run via FileSystemOutputStore to the repo-root .agenteval/ workspace (resolved via the same *.sln/*.slnx/.git/ walk-up agenteval init uses). Manifest, scenarios, summary, and compliance evidence land there. Sidecar HTML / PDF / bare JSON stay project-local under samples/AgentEval.Samples/output/.
  • Compliance reporters (GDPR, EU AI Act, OWASP, MITRE) invoked end-to-end on the matching samples with full audit-chain anchoring (sourceRun.manifestHash matches the canonical manifest's contentHash exactly).
  • No stubs, no hardcoded responses — Performance uses a real Azure agent (not EchoAgent); Agentic invokes the agent live (not a hardcoded response); GDPR + EU AI Act probe the agent per scenario (not a single shared response). OWASP + MITRE already exercised real agents.
  • H2 Performance is metric-only (latency / throughput / cost). Every other running sample (H3 onward) uses a real Azure-backed agent and a real LLM judge.
  • SamplePreset toggle — AGENTEVAL_SAMPLES_PRESET=smoke|standard|audit-grade (or --preset <value> CLI arg) scales runtime from cents to audit-grade. Default: smoke.
  • Mission Control integration — discovers .agenteval/ via --workspace flag (now honoured on bare dotnet run --project too); surfaces subjects + runs + compliance summaries.
  • Open-after-save prompt — [h]/[j]/[p]/[n] after each sample writes its reports; cross-platform Process.Start(UseShellExecute=true); honours AGENTEVAL_SAMPLES_NONINTERACTIVE=1 and redirected stdin (skips cleanly in CI).
  • H10 Report Browser sample — interactive past-run browser; one-keystroke open of any JSON / HTML / PDF.
  • H9 Memory sample — Shape-B bridging over the canonical MemoryBenchmarkRunner; MemoryBenchmarkResult synthesised into an EvalResult tree, native shape preserved in report-native.json.
  • H8 LongMemEval sample — promoted from metadata-only walkthrough to a preset-driven (Smoke / Standard / AuditGrade) running sample against the real longmemeval_s_cleaned.json dataset; friendly download-instructions box when the dataset is missing.

Breaking changes

The bulk of v0.10.1 is purely additive on top of v0.10.0-beta (new renderers, new sample suite, new canonical-store wiring). The "real-data-only" LongMemEval shift, however, removes one previously-public API and tightens dataset-path resolution. NuGet consumers depending on these surfaces will need to migrate:

  1. LongMemEvalDataLoader.LoadEmbedded(...) removed. The bundled "inspired by LongMemEval" subset (10 entries, partial schema) was a hand-authored approximation that produced misleading scores — both the static method and the underlying embedded resource are gone. Migration: replace LongMemEvalDataLoader.LoadEmbedded(options) with LongMemEvalDataLoader.LoadResolved(options) and ensure the real longmemeval_s_cleaned.json is reachable via canonical local path (<workspace-root>/src/AgentEval.Memory/Data/longmemeval/) or the LONGMEMEVAL_DATASET_PATH env var. Catch LongMemEvalDatasetNotFoundException for friendly download-instructions UX (see samples/AgentEval.Samples/Benchmarks/08_LongMemEvalBenchmark.cs for the pattern).
  2. LongMemEvalDataLoader.ResolveDatasetPath(...) tightened semantics. When a non-whitespace explicitPath argument or the LONGMEMEVAL_DATASET_PATH env var is supplied but the file does NOT exist on disk, the method now throws LongMemEvalDatasetNotFoundException instead of silently falling through to the env var / canonical local path. The previous behaviour could silently run a benchmark against a different dataset than the caller asked for — a misleading-results bug for typos in DatasetPath or the Full() env-var path. Fall-through to the canonical local path now only applies when neither explicit nor env-var is supplied. Migration: validate File.Exists at the call site before invoking, or catch LongMemEvalDatasetNotFoundException and surface the typo to the user.

Install

<PackageReference Include="AgentEval" Version="0.10.1-beta" />

Known issues — deferred to v0.10.2+

Tracked in CHANGELOG.md [Unreleased]:

  • Pre-existing LLM non-determinism flake on SafetyPolicyTests.CancellationRequest_ShouldConfirmBeforeCancelling.
  • Missing docs/redteam/owasp.md (referenced from OwaspBenchmarkRegistration.docLinkUrl).
  • README.md benchmark-table sweep + docs/benchmarks.md update.
  • Agentic-safety + GDPR/EuAiAct domain-pack registry surface flags.

See CHANGELOG.md for the full v0.10.1-beta entry (Added / Changed / Removed / Breaking sections).