Skip to content

v0.5.1 - Reports and package publishing

Choose a tag to compare

@guybass guybass released this 14 Sep 11:22

Agent Eval Flow turns agent execution evidence into reports that help explain failures and compare improvements.

Available on PyPI:

python -m pip install agent-eval-flow==0.5.1

This release adds a live OpenSRE/OpenKritt report gallery, screenshot previews, task-and-fix walkthroughs, repository visuals, issue forms and package discovery metadata. It also adds a Trusted Publishing workflow that verifies the release version and all six CI jobs before uploading to PyPI.

The evaluation behavior is unchanged from 0.5.0. This remains a developer preview: the two case studies are small, controlled experiments, not general agent-reliability or security-accuracy benchmarks.

Validation: strict wheel/source metadata checks, all 90 frozen fixture hashes preserved, and all six Ubuntu/Windows and Python 3.11–3.13 CI jobs passed. A fresh installation from PyPI loaded, validated and rendered both saved-result formats. The attached wheel and source archive match PyPI's SHA-256 hashes; checksums are included.

Only curated example reports are included. Raw local captures and unpublished social-media drafts are excluded.

See the release guide and offline quickstart.