Skip to content

v0.2.2

Choose a tag to compare

@janit janit released this 20 Sep 20:41
· 6 commits to main since this release

The planner's reasoning is kept, every item gets rated, and the two empty
answers stop looking alike.

gator auto plan ran against a real model for the first time on 2026-09-20 β€”
every test until then used a shell stub. It selected correctly in 2 of 4 runs,
and worse, the failures could not be diagnosed: the raw response was thrown
away the instant it was parsed. Reproducing a run by hand showed the model had
reasoned well and the tool had discarded it.

After the changes below, on the same source and the same model: 3 of 4, with
all seven items considered every run and every rejection explained.

The response is kept

Always at .gator/auto/last-response.txt, written before validation β€” a
response that fails to validate is the one worth reading β€” and under the plan's
own id once a plan exists. 0600, like the rest of the store; it can quote the
source.

Rate every item; the code chooses

The prompt used to say "select one unit of work", and the model reasonably
returned one candidate. So rejected was always empty and the
rejected-alternatives explanation was lost, even though the model had reasoned
about everything in prose first.

It now asks for every item the source contains, rated, and says plainly that
selection happens in code from those ratings. The model's judgement did not
change; the tool stopped throwing most of it away.

Three outcomes, not two

selected, none_eligible (items were considered and each turned down β€”
considered says how many) and no_items (the planner found nothing at all).
The last two printed the same line before, which made a planner malfunction
indistinguishable from a correct abstention. no_items now points at the
response so it can be judged.

Validation is per candidate

Asking for every item means being sent items that cannot be built β€” an
undesigned task names no files, so its scope is empty. Refusing the whole plan
over one of those destroyed three otherwise-correct plans in testing.

A candidate that fails the contract is now reported with its reason and
excluded from selection. Nothing unvalidated can be selected, which is the
property that mattered; a malformed response still refuses wholesale.

Note

auto plan now asks for more and therefore takes longer β€” seven rated
candidates instead of one. Raise GATOR_PLANNER_TIMEOUT if your planner is
slow.

gator auto run still does not exist.