v0.10.1-beta — Samples Consolidation + Generic Renderers + Canonical Store
Adds a uniform
IEvalResultRenderercontract (HTML + PDF) inAgentEval.Abstractions, with implementations inAgentEval.Core(HTML) and the newAgentEval.Rendering.Pdfproject. Consolidates the per-family*.Demoprojects into one focusedsamples/AgentEval.Samples/Benchmarks/suite — 10 examples wired as menu group H. Every running sample writes the canonical run throughFileSystemOutputStoreto the repo-root.agenteval/workspace so Mission Control +agenteval doctorsee it, plus a sidecaroutput/{family}/run-{utc}-{suffix}/for direct human consumption of HTML + PDF.
End-to-end verified against a real Azure agent on gpt-5-chat before merge; full audit-chain integrity (manifest contentHash ↔ HTML/PDF footer ↔ compliance evidence sourceRun.manifestHash) confirmed.
Highlights
IEvalResultRendereruniform rendering contract inAgentEval.Abstractions—RenderAsync(EvalResult, options) → byte[]. Framing metadata (subject, run id, audit hash, AgentEval version) flows throughEvalResultRenderOptions.HtmlEvalResultRenderer(Core) — self-contained HTML, inline CSS,<details>collapsibles, XSS-safe viaWebUtility.HtmlEncode, severity- and label-coded badges (threshold-fail composites no longer render green), honestNOT TESTEDfor skipped leaves on both leaves AND the cover banner.AgentEval.Rendering.Pdfnew project withPdfEvalResultRenderer— QuestPDF-backed; cover page + component summary + per-leaf detail pages + audit-chain appendix. Embedded into the umbrella NuGet viaPrivateAssets="all".samples/AgentEval.Samples/Benchmarks/sample suite — 10 focused examples in menu group H: Registry Discovery, Performance, Agentic, GDPR, EU AI Act, OWASP, MITRE, LongMemEval, Memory, Report Browser.- Canonical store wiring — every running sample writes a canonical run via
FileSystemOutputStoreto the repo-root.agenteval/workspace (resolved via the same*.sln/*.slnx/.git/walk-upagenteval inituses). Manifest, scenarios, summary, and compliance evidence land there. Sidecar HTML / PDF / bare JSON stay project-local undersamples/AgentEval.Samples/output/. - Compliance reporters (GDPR, EU AI Act, OWASP, MITRE) invoked end-to-end on the matching samples with full audit-chain anchoring (
sourceRun.manifestHashmatches the canonical manifest'scontentHashexactly). - No stubs, no hardcoded responses — Performance uses a real Azure agent (not
EchoAgent); Agentic invokes the agent live (not a hardcoded response); GDPR + EU AI Act probe the agent per scenario (not a single shared response). OWASP + MITRE already exercised real agents. - H2 Performance is metric-only (latency / throughput / cost). Every other running sample (H3 onward) uses a real Azure-backed agent and a real LLM judge.
SamplePresettoggle —AGENTEVAL_SAMPLES_PRESET=smoke|standard|audit-grade(or--preset <value>CLI arg) scales runtime from cents to audit-grade. Default:smoke.- Mission Control integration — discovers
.agenteval/via--workspaceflag (now honoured on baredotnet run --projecttoo); surfaces subjects + runs + compliance summaries. - Open-after-save prompt —
[h]/[j]/[p]/[n]after each sample writes its reports; cross-platformProcess.Start(UseShellExecute=true); honoursAGENTEVAL_SAMPLES_NONINTERACTIVE=1and redirected stdin (skips cleanly in CI). - H10 Report Browser sample — interactive past-run browser; one-keystroke open of any JSON / HTML / PDF.
- H9 Memory sample — Shape-B bridging over the canonical
MemoryBenchmarkRunner;MemoryBenchmarkResultsynthesised into anEvalResulttree, native shape preserved inreport-native.json. - H8 LongMemEval sample — promoted from metadata-only walkthrough to a preset-driven (Smoke / Standard / AuditGrade) running sample against the real
longmemeval_s_cleaned.jsondataset; friendly download-instructions box when the dataset is missing.
Breaking changes
The bulk of v0.10.1 is purely additive on top of v0.10.0-beta (new renderers, new sample suite, new canonical-store wiring). The "real-data-only" LongMemEval shift, however, removes one previously-public API and tightens dataset-path resolution. NuGet consumers depending on these surfaces will need to migrate:
LongMemEvalDataLoader.LoadEmbedded(...)removed. The bundled "inspired by LongMemEval" subset (10 entries, partial schema) was a hand-authored approximation that produced misleading scores — both the static method and the underlying embedded resource are gone. Migration: replaceLongMemEvalDataLoader.LoadEmbedded(options)withLongMemEvalDataLoader.LoadResolved(options)and ensure the reallongmemeval_s_cleaned.jsonis reachable via canonical local path (<workspace-root>/src/AgentEval.Memory/Data/longmemeval/) or theLONGMEMEVAL_DATASET_PATHenv var. CatchLongMemEvalDatasetNotFoundExceptionfor friendly download-instructions UX (seesamples/AgentEval.Samples/Benchmarks/08_LongMemEvalBenchmark.csfor the pattern).LongMemEvalDataLoader.ResolveDatasetPath(...)tightened semantics. When a non-whitespaceexplicitPathargument or theLONGMEMEVAL_DATASET_PATHenv var is supplied but the file does NOT exist on disk, the method now throwsLongMemEvalDatasetNotFoundExceptioninstead of silently falling through to the env var / canonical local path. The previous behaviour could silently run a benchmark against a different dataset than the caller asked for — a misleading-results bug for typos inDatasetPathor theFull()env-var path. Fall-through to the canonical local path now only applies when neither explicit nor env-var is supplied. Migration: validateFile.Existsat the call site before invoking, or catchLongMemEvalDatasetNotFoundExceptionand surface the typo to the user.
Install
<PackageReference Include="AgentEval" Version="0.10.1-beta" />Known issues — deferred to v0.10.2+
Tracked in CHANGELOG.md [Unreleased]:
- Pre-existing LLM non-determinism flake on
SafetyPolicyTests.CancellationRequest_ShouldConfirmBeforeCancelling. - Missing
docs/redteam/owasp.md(referenced fromOwaspBenchmarkRegistration.docLinkUrl). README.mdbenchmark-table sweep +docs/benchmarks.mdupdate.- Agentic-safety + GDPR/EuAiAct domain-pack registry surface flags.
See CHANGELOG.md for the full v0.10.1-beta entry (Added / Changed / Removed / Breaking sections).