Skip to content

v0.0.25

Choose a tag to compare

@aviggiano aviggiano released this 10 Sep 11:22
0577d17

v0.0.25 makes new runs use best-effort completion and gives operators a clear report outcome when tasks fail. Final report content remains agent-written.

  • Keep useful work moving. Independent tasks continue after ordinary failures. Review uses successful, checked inputs, and incomplete coverage produces a PARTIAL report when the report agent succeeds.
  • Know when reporting is unavailable. status --watch --json exposes whether execution ended and the report's availability, completion, verification, and paths. Failed reporting keeps the saved results without creating a replacement report.
  • Choose the required checks. run --require-complete requires complete execution coverage. report --require-verified separately requires report verification. Neither option adds retries.

Breaking changes

  • [runtime] [config] Adds default best-effort completion, agent-written PARTIAL reporting, and report availability in the existing status command. Local report verification is optional; strict bundles, benchmark scoring, and public publication retain their verification requirements. The sealed resolved-config contract advances to v4. This release applies to new runs; migration and resumption of older runs under the new contract are unsupported. Custom topologies retain their declared failure policies, and configured attempts and time limits still apply. Updates the TOML parser to 1.7.1 to address GHSA-7w5x-hrqm-74c2. Thanks @aviggiano! (#1120)

Improvements

  • [benchmarks] Refreshes published benchmark history and charts through the automatic publisher. Direct commit: 46ae7dd2.

Bug fixes

  • [evals] Fixes the required-backend preflight test so clean CI runners verify the intended rejection without depending on reference caches. Thanks @aviggiano! (#1122)

Published without waiting for the final PR CI rerun or CI on the release commit, as authorized by the maintainer. Local validation includes all 187 CLI tests and all 450 evaluation tests passing.

Full changelog: v0.0.24...v0.0.25