Skip to content

v0.2.1 — honest failure edges for long-running coding sessions

Latest

Choose a tag to compare

@protostatis protostatis released this 03 Sep 01:49
· 46 commits to main since this release

Patch release on top of v0.2.0. Full series notes: the 0.2.x line ships the long-running coding-session vertical slice (task contracts, user-owned acceptance checks, verified/failed/not-verified states, bounded repair, cache-aware context accounting) plus the evaluation layer (real-engine suite runner, scripted mock transport, verifier-assisted opt-in policy, campaign audit).

What v0.2.1 fixes (from the post-release advisor review):

  • A failed loop-breaker model request is now recorded as a provider error and graded infrastructure_error, never a fabricated pass/fail.
  • A repair response that used tools no longer ends the turn silently: one-shot mode gets a deterministic, clearly-labeled completion notice.
  • Acknowledged host-mode evaluation runs can pass the audit (containment requirement now matches the mode actually executed; execution_mode recorded per attempt).
  • /resume no longer restores transient context-prefix state; cumulative counters still survive.
  • /cwd persists the acceptance-contract rebind, and resuming in a different working directory invalidates stale verified status.
  • One-shot mode exits nonzero when a provider failure ends the turn.
  • Acceptance evidence carries a checked_at timestamp; README claims tightened.

136 tests pass; self-test passes; the existing 40-attempt evaluation records still audit VALID.