Skip to content

Releases: janit/gator

v0.4.1

Choose a tag to compare

@janit janit released this 27 Sep 08:56

The classifier now recommends granite-4.2-8b: 32/32 on the labelled and borderline sets at ~65 ms per call, against ~700 ms for Qwen3.8-27B. Classifier task text is capped at 4 KB, down from 16 KB, which keeps a call inside the default 5 s timeout on slower GPUs. Stored plans are not marked stale by the cap change.

v0.4.0

Choose a tag to compare

@janit janit released this 24 Sep 14:45

Merges are verified as they will land

  • The integration candidate is verified, not just the branch. When the target branch has moved since a unit started β€” another unit merged first, say β€” the merge is built in a scratch worktree, the unit's verifier runs on that exact tree, and the target fast-forwards to it only when green. Two units that each pass alone but fail together now report integration_failed instead of merging a red tree.
  • A unit merges only into the branch it was fed from. Switch branches before finalising and the unit is held, then merges once you switch back.

Behaviour changes

  • --scope accepts file paths and dir/** directories. Other wildcards (src/*.ts) are refused at feed time β€” they were previously accepted and silently not enforced. "**" still declines scoping.
  • A unit reported once but not merged (failed, conflicted, unverified, out of scope, integration failed) stays listed by gator status / gator wait until its branch is merged or deleted.
  • feed, drain and the finalise pass are serialised by a per-repository flock, so concurrent callers cannot race on the budget, the queue or merges. flock (util-linux) is now a runtime requirement.

Fixes

  • An interrupted queue drain can no longer launch the same unit twice.
  • gator sched status reports reserved only for backends configured to honour reservations.
  • gator auto plan --from with no value is refused with [missing_value] instead of a Python traceback.
  • auto plan prunes worktrees left registered by an interrupted baseline check.
  • The test suite runs under a per-run temporary directory and removes it afterwards.

Docs

  • docs/auto.md reason codes match the code.
  • GATOR_VERIFY_TIMEOUT and GATOR_POLL are documented.
  • Keep API keys out of worker_cmd: the command appears in process arguments, which other local users can read.

v0.3.0

Choose a tag to compare

@janit janit released this 21 Sep 19:51

Heterogeneous local GPU scheduling. On a machine with more than one local
GPU, gator now decides where a unit may run, separately from which model
should run it.

Preference may move ordinary work; eligibility may not move heavy work.

Resource classes and eligibility

A resources file declares the backends and which of them each class may use.
Eligibility is a hard filter applied before any scoring, so no weighting can
talk a heavy unit onto the smaller card. If the card it needs is busy, the
unit queues; if the card is down, the unit waits. A property test over 500
random weightings asserts this.

Leases and a queue with no daemon

  • Backend leases are flocked JSON documents, replaced atomically. A worker
    that crashes has its lease reaped by process liveness. A lease file that will
    not parse raises an error instead of being read as "no leases held".
  • The queue is drained when something is already happening (feed,
    status/wait, and the release of the blocking lease), so nothing resident
    has to be supervised.
  • gator sched inspects and controls backends, e.g. gator sched disable.
  • Once a queued unit's only eligible backend has been down for longer than
    GATOR_BACKEND_TIMEOUT (default 1800 s), the unit is reported as
    blocked_backend. It stays queued for that backend and is never re-routed.
    Draining the queue now respects GATOR_PROBE_HEALTH, so a queued unit is no
    longer launched onto a dead backend.

Task-difficulty classifier (opt-in)

When nothing the operator configured decides a unit's class, gator can ask a
model how hard the task is. It reads the model's logprobs for two options
(heavy vs standard) from a single decode step and uses none of the text the
model writes. Units start at heavy and are demoted only above a confidence
threshold. Every failure (timeout, unreachable endpoint, unparseable reply)
keeps the unit heavy, and the record says why. Every operator setting
outranks the classifier. The endpoint passes the same trust gate as the rest of
the resources file.

gator auto plan records the classifier's recommendation on each selectable
candidate inside the hashed plan, and changing the classifier settings marks a
stored plan stale_policy.

Unchanged

Verification, merge-on-green, worktree isolation, the record and budgets. The
scheduler decides where a unit runs, not whether its result is any good.

See docs/scheduling.md. Upgrade with deno task install.

v0.2.4

Choose a tag to compare

@janit janit released this 21 Sep 07:42

Security fix. An unusable candidate's id could forge a rejection line in
the terminal.

When per-candidate validation landed in v0.2.2, a candidate that fails the
contract began being reported rather than refusing the whole plan. The
rejection record carries candidate.get("id") β€” a value that by construction
has been through no validation, that being exactly why it is in the error
handler β€” and the renderer prints it.

A planner returning an id of:

bad<ESC>[2K<CR>  rejected fake-id: looks fine

erases the rejection line gator just printed and writes its own in its place.
The model whose output is being judged can forge the judgement, for whoever
reads the terminal.

The existing guard in validate_candidates was never the problem and still
works: a valid candidate's free text has been scanned for control characters
since v0.2.2. The reporting path added beside it had no such check.

The fix

safe_text() applied at the data boundary, not at the print site. The
manifest is stored and re-rendered by other commands later, so escaping only
when printing would leave a delayed version of the same attack sitting on
disk.

Four tests, wider than the coverage that existed before: both refusal paths
including acceptance-criteria scanning, the id reaching a terminal, and no raw
control byte in the stored manifest.

Affects

v0.2.2 and v0.2.3, wherever gator auto plan runs against a planner that can
be induced to return a hostile candidate id. Upgrade with deno task install.

v0.2.3

Choose a tag to compare

@janit janit released this 21 Sep 06:13

--scope is now enforced against the diff, rather than requested in the
prompt.

Until now the scope you gave a unit was prose interpolated into the worker's
instructions, and nothing ever compared it to what the worker actually did. A
unit told to touch src/** could write anywhere in its worktree and merge. The
reproduction has been in the test suite since a security review on 2026-09-20,
pinned as a description of what the tool did; it now describes what it
refuses to do.

UNIT "touches-readme" β†’ out_of_scope  [0s, 1 commit(s), branch gator/touches-readme]
NOT merged: the unit changed files outside the scope it was given.
declared scope: src/**
outside it:
  README.md
--- worktree kept at .gator/worktrees/touches-readme ---

What is checked

The whole base-to-result diff: both endpoints of a rename, deletions, new
files, mode changes and symlinks. Each is a way to move work out of the
declared area, and a check that looked only at added files would miss every one
of them. Git's output is read NUL-delimited, so a path cannot hide inside a
newline.

A violation holds the merge, keeps the worktree for inspection, and names the
offending paths. The unit is atomic β€” the in-scope half is not merged either.

One matcher

Validation and enforcement share it. If the rules a planner must satisfy
differed from the rules a worker is held to, the scope a model may declare
would drift from the scope it is judged against.

The single deliberate difference: you may write --scope "**" to decline
scoping, because that is a choice a person can make and be asked about. The
planner cannot β€” a whole-repository scope proposed by the thing being judged is
not a scope.

Precedence

Following the specification: the worker's own failure outranks scope, and scope
outranks a failing verifier. A unit that wrote where it should not is not
redeemed by a green build, and "it wrote outside its scope" is the more
actionable thing to tell you.

Upgrading

deno task install. If you have been relying on broad scopes implicitly, units
that stray will now be held rather than merged β€” widen the scope deliberately,
or use "**".

v0.2.2

Choose a tag to compare

@janit janit released this 20 Sep 20:41

The planner's reasoning is kept, every item gets rated, and the two empty
answers stop looking alike.

gator auto plan ran against a real model for the first time on 2026-09-20 β€”
every test until then used a shell stub. It selected correctly in 2 of 4 runs,
and worse, the failures could not be diagnosed: the raw response was thrown
away the instant it was parsed. Reproducing a run by hand showed the model had
reasoned well and the tool had discarded it.

After the changes below, on the same source and the same model: 3 of 4, with
all seven items considered every run and every rejection explained.

The response is kept

Always at .gator/auto/last-response.txt, written before validation β€” a
response that fails to validate is the one worth reading β€” and under the plan's
own id once a plan exists. 0600, like the rest of the store; it can quote the
source.

Rate every item; the code chooses

The prompt used to say "select one unit of work", and the model reasonably
returned one candidate. So rejected was always empty and the
rejected-alternatives explanation was lost, even though the model had reasoned
about everything in prose first.

It now asks for every item the source contains, rated, and says plainly that
selection happens in code from those ratings. The model's judgement did not
change; the tool stopped throwing most of it away.

Three outcomes, not two

selected, none_eligible (items were considered and each turned down β€”
considered says how many) and no_items (the planner found nothing at all).
The last two printed the same line before, which made a planner malfunction
indistinguishable from a correct abstention. no_items now points at the
response so it can be judged.

Validation is per candidate

Asking for every item means being sent items that cannot be built β€” an
undesigned task names no files, so its scope is empty. Refusing the whole plan
over one of those destroyed three otherwise-correct plans in testing.

A candidate that fails the contract is now reported with its reason and
excluded from selection. Nothing unvalidated can be selected, which is the
property that mattered; a malformed response still refuses wholesale.

Note

auto plan now asks for more and therefore takes longer β€” seven rated
candidates instead of one. Raise GATOR_PLANNER_TIMEOUT if your planner is
slow.

gator auto run still does not exist.

v0.2.1

Choose a tag to compare

@janit janit released this 20 Sep 19:39

gator auto now reads the roles file, which is how roles were always meant
to work.

The controller resolved roles from the environment only, so the documented
mechanism β€” a roles file, no model id in the code, no shell profile edited β€”
did not apply to auto. Worse, it failed opaquely: with no model resolved the
placeholders went through unsubstituted and the host answered with its own
Unknown provider "%PROVIDER%".

# ~/.config/gator/roles
planner = yeti/DeepSeek-V4-Flash    # used by `gator auto plan`
heavy   = yeti/DeepSeek-V4-Flash    # fallback, and what `feed` uses

Precedence lives in one place. The shell script already implemented it β€”
environment, then the repository's file when trusted, then yours β€” and now
resolves planner, heavy and planner_cmd before handing over, so the
controller receives answers rather than re-deriving the rules in a second
language. A missing planner role refuses with planner_role_not_configured and
says what to set.

Model references from a roles file go through the same validator feed uses,
since auto substitutes them into a command too.

No test caught this: every planner in the suite is a stub. That is a gap in the
tests as much as a bug in the code.

Upgrading from v0.2.0 is deno task install; nothing else changes.

v0.2.0

Choose a tag to compare

@janit janit released this 20 Sep 19:39

Recommendation-only auto planning, and the repository stops being trusted config.

gator auto plan reads one committed file you name and recommends a single unit of
work from it. It recommends only β€” no worker starts, nothing is merged.

gator auto plan  --from SPEC.md     # recommend one unit
gator auto show  --plan <plan-id>   # inspect it, and whether it still applies
gator auto plans                    # what has been planned here

The result is canonical, hashed JSON bound to this checkout, the branch, its HEAD,
the source's committed blob and the selection policy. Change any of them and the
plan stops applying; show names which one. "No suitable work" is a normal,
successful answer
and exits 0.

Selection is deterministic and explainable β€” a rubric of benefit, clarity,
boundedness and risk, with every rejection carrying its reason. The planner runs
with no tools at all, and the source it reads is fenced as untrusted data.

The repository is no longer trusted configuration

.gator/ lives inside the repository, and a repository can commit its own β€”
.git/info/exclude suppresses only untracked files. So a clone could arrive
carrying a roles file naming the command to run, or a verify file that would
be executed. Both are now ignored unless you opt in:

GATOR_TRUST_REPO_CONFIG=1 gator ...   # this repository's .gator/ is mine

Without it they are skipped with a note on stderr. Set it only for repositories
you wrote. Model references are validated before substitution, so one cannot
carry a subshell.

Also

  • Structured unit records: base SHA, target ref, the exact verifier that ran and
    the commit it verified. A unit that exits nonzero without committing is now
    failed, not empty β€” the failure cause outranks the absence of output.
  • gator status --json and --no-merge are pure reads: no merge, no worktree
    removal, no reported marker.
  • Planning is rate limited β€” cooldown, budget, one at a time β€” because each call
    costs a model request and a full verifier run.
  • Planner output is streamed against a hard cap rather than buffered and then
    measured. Control characters in planner strings are refused, so a model cannot
    rewrite the terminal output that judges it.
  • The store is 0700 with 0600 files: it holds prompts, model output and
    verifier logs.

Not in this release

gator auto run does not exist. Execution arrives once the safety tests for it
pass; until then auto plan produces something for a human to read and act on.
Worktrees protect integration, not the machine β€” see the security model in the
README.

v0.1.0

Choose a tag to compare

@janit janit released this 19 Sep 12:52

Feed the gator. Hand a chunk of work to a bigger model and it chews on it in
an isolated git worktree on its own branch β€” merged back only when it builds and
tests clean.

gator feed --title "port the tokenizer" --role heavy \
           --scope "src/lexer.ts,src/lexer_test.ts" --task @task.md
gator wait

One SKILL.md and two shell scripts. No plugin, no host API: any agent that reads
skills/<name>/SKILL.md and can run a shell command can use it β€” OpenCode, Pi,
Claude Code.

Why a script and not a plugin tool

Delegation fails at one step: getting the model to choose it. Measured against a
603-line build brief with Qwen3.8-27B as the chat model, a delegate tool reached
through a plugin's code mode was called in 0 of 5 runs; the same model reached
for this script in 5 of 7 (Fisher exact, two-sided p = 0.0278). In one of those
runs the model made 179 shell calls and 2 execute calls. shell is the tool
these models live in, so that is where feeding belongs.

The same benchmark at the API level: with a file-reading tool available, neither
granite-4.2-8b nor Qwen3.8-27B delegated β€” 0/9 each. Remove read and Qwen goes
9/9. Model size was never the constraint.

Merge only when green

Each chunk is built and tested inside its own worktree before anything is
merged, so a broken one costs a merge that never happened rather than a rollback
of your work. In the same study a delegated implementation scored 90/301 as
delivered and 299/301 after fixing two characters in an import path β€” 209
points lost because nothing ran the build. A chunk that does not build comes back
unverified, keeps its worktree, and hands you the build output.

The verify command is detected from the project (deno task build && deno task test, npm test, cargo test), or set in .gator/verify or GATOR_VERIFY.

What else is in it

  • Limits enforced in the script, not asked of a model: a minimum chunk size, scope
    required, dedup by task hash, a feeding budget, digestion time between feedings,
    and a cap on how many chunks it chews at once.
  • Chunks are chewed detached, so one outlives the tool call that fed it and
    survives the session being interrupted.
  • Roles live in ~/.config/gator/roles, so no model id is in the code and no shell
    profile has to be edited.
  • Statuses that refuse to flatter: empty when it spat the chunk out, unverified
    when it does not build, ready (held) when your tree is dirty.
  • deno task test stubs the worker, so every limit, every status and the
    verification gate run in about seven seconds without touching a model.

docs/evidence.md has the measurements behind each decision, including what they
do not prove: one fleet, two quantised local models, one workload.