Repository navigation
Releases: guybass/agent-eval-flow
Release list
v0.6.0 - Toolscore diagnostics and native-agent case studies
Agent Eval Flow 0.6.0 adds deterministic tool-call evaluation and offline case studies, and fixes two failures encountered in native-agent captures.
Changes
- Optional Toolscore 1.9.0 integration with eight diagnostics, including identical-call rate, source-linked receipts, and portable tools reports alongside the general report. Correct empty traces receive a perfect native score; incomplete traces stay unknown.
- Claude stream projection accepts system and result records without discarding the retained stream. Unsupported assistant/user message content still fails explicitly. (#5)
- Completed grading survives a backwards host-clock adjustment. (#4)
- Offline GPT Researcher, DeerFlow, and MCP memory examples import retained evidence and keep outcomes separate from tool diagnostics. MCP importers preserve tool errors and use the bundled scenario file. (#1, #2, #3)
- DeerFlow summaries retain observed request counts when the optional Toolscore package is absent, with the missing scorer reported explicitly.
- Security-triage research documentation records uncertainty loss, development regressions, failed attempts, and evidence limitations. (#8)
- Publishing verifies all twelve offline and Toolscore CI jobs across Windows/Linux and Python 3.11, 3.12, and 3.13. Regression tests reject missing, skipped, duplicated, or failed job coverage.
Upgrade
python -m pip install --upgrade "agent-eval-flow[toolscore]==0.6.0"
Toolscore is optional. The base installation and rendering saved reports do not require it. Existing 1.8.1 receipts keep their original scores. For new comparisons, regrade both candidates with the same evaluator: one-to-one argument pairing and correct no-tool scoring can change existing scores.
Validation and scope
The release passes the offline and Toolscore CI matrices, package integrity and metadata checks, and offline report examples. All 85 frozen acceptance files retain their original hashes. The included studies are documented development evidence, not a new untouched holdout or a general reliability estimate; this release performs no new live-agent evaluation.
The contract-v2 and Claude subscription-auth proposal (#7) is deferred pending compatibility, report-count, and isolation validation. Those experimental APIs are not included in 0.6.0; the original experimental branch and frozen study remain unchanged.
Merged contributions: #2, #3, #4, #5, #6, and #8. The existing Toolscore integration and GPT Researcher changes are included in this first PyPI release since 0.5.1.
v0.5.1 - Reports and package publishing
Agent Eval Flow turns agent execution evidence into reports that help explain failures and compare improvements.
Available on PyPI:
python -m pip install agent-eval-flow==0.5.1This release adds a live OpenSRE/OpenKritt report gallery, screenshot previews, task-and-fix walkthroughs, repository visuals, issue forms and package discovery metadata. It also adds a Trusted Publishing workflow that verifies the release version and all six CI jobs before uploading to PyPI.
The evaluation behavior is unchanged from 0.5.0. This remains a developer preview: the two case studies are small, controlled experiments, not general agent-reliability or security-accuracy benchmarks.
Validation: strict wheel/source metadata checks, all 90 frozen fixture hashes preserved, and all six Ubuntu/Windows and Python 3.11–3.13 CI jobs passed. A fresh installation from PyPI loaded, validated and rendered both saved-result formats. The attached wheel and source archive match PyPI's SHA-256 hashes; checksums are included.
Only curated example reports are included. Raw local captures and unpublished social-media drafts are excluded.
See the release guide and offline quickstart.
v0.5.0 - Developer preview
Version 0.5.0 — developer preview
Agent Eval Flow evaluates complete agent systems through their existing runtimes
or retained execution logs. This initial public release is intended for developers
building evaluation workflows and inspecting failures.
Included
- Typed studies, candidates, tasks, execution evidence and evaluation results.
- Import and regrade saved runs without invoking the target agent again.
- Parallel configuration inspection and behavioral evaluation with explicit
decision policies. - Optional runtime observations and deterministic checks for instructions,
tools, model settings, loops, memory and environment state. - HTML reports that retain measurement status, provenance and expected versus
observed values. - Runnable offline examples and a curated OpenSRE/OpenKritt report
gallery.
Verification
Final local release validation on 2026-09-13 passed 434 tests, with 21 expected
skips: twenty unselected live-profile cases and one POSIX-only process-group
test on Windows. Both offline examples passed, and all 90 frozen acceptance
and fixture hashes were unchanged. No live model profiles were selected.
The source distribution and wheel built successfully; package contents matched
the implementation and preserved the fixture hashes. The installed wheel loaded
and rendered both saved result formats and exposed the runtime-evidence modules.
All six GitHub Actions jobs passed on Ubuntu and Windows with Python 3.11,
3.12 and 3.13. Verified CI run.
The timeout regression now allows the fixture's normal startup interval and
checks that the report process actually timed out, retained stdout/stderr,
stopped its direct child and cleaned up its staging directory. It no longer
assumes a new Python interpreter starts within 300 ms.
Fresh-checkout CI also checks pinned workflow bytes under Git's different
line-ending modes and reads native artifact paths without dropping Windows
drive letters or interpreting literal path characters as URL syntax.
Scope
This is an initial implementation. Native adapters require version-specific
runtime setup. The shared runtime-observation contract does not automatically
interpret arbitrary logs; collectors must explicitly capture and map observations.
The two report case studies use small synthetic development tasks with adaptive
changes. They demonstrate specific system improvements, not general reliability,
security accuracy or isolated changes in model reasoning. Their raw local
captures, experimental harnesses and social-media drafts are not distributed.
The runnable archive example uses a separately licensed, pinned upstream
trajectory. Retained upstream test fixtures are intentionally included as test
inputs; see third-party notices.
Try it
Start with the installation instructions and offline example. The attached wheel and source archive have SHA-256 checksums in SHA256SUMS.
A project license has not yet been selected; third-party fixtures retain their documented terms.