Skip to content

1.17.0 - the stopping decision, checked exhaustively

Choose a tag to compare

@ridelink0 ridelink0 released this 08 Sep 22:28
· 87 commits to main since this release

1.17.0 - the stopping decision, checked exhaustively

The complaint that started all of this was specific: under Codex it "doesn't
stop at a good stopping point", and it "just makes a plan". Both are failures
of one decision - what the budget line tells the agent to DO - and that
decision now has five inputs: how full the binding window is, whether it is
scoped to a model or shared by the account, which host is reading, whether a
relay is armed, and whether a cheaper lever exists.

Five inputs is too many to spot-check, and spot-checking had already missed
one. test/stopping.test.js walks all 160 combinations and asserts the rules
the whole matrix must satisfy: never tell a spent shared window it still has
budget; always offer the switch on a spent model-scoped one; never say stop
and carry on in the same line; never ask for a handoff below the wall; never
name a Claude Code command to Codex; never leave a stop instruction standing
with a relay armed; and never emit an empty line or leak a value into one.

It failed on its first run. The effort sentence built its command from the
process host rather than the host the line was being written for, so a line
composed for Codex could carry /effort - a command that does not exist there.
Right in production by accident, wrong in principle, and now the host is
passed in rather than read from global state.

That is two real bugs from two walks of this decision: the previous one had a
spent shared window claiming it was "not out of budget" because a cheaper
effort existed. Both are the same underlying mistake - not distinguishing
which window binds and what a lever actually does to it - which is exactly why
it is a matrix now instead of a habit of checking carefully.

630 tests.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com