Repository navigation
Releases: leanzero-srl/goose-local-edition
Release list
3.0.108 Cut-off replies continue; Forge charges malformed manifests
3.0.108 — cut-off replies are continued, and Forge charges malformed manifests instead of crashing
What happened
Three problems surfaced on 2026-10-07:
- aion-labs/aion-3.5's Gauntlet run kept stalling. Twice in a row the model spent all 32,768 output tokens
allowed by its only host on reasoning, with no text and no tool call. The final chunk of each stream carried
nothing, so goose never marked the reply as cut off. The agent treated each reply as empty and sent the same
prompt again. After three retries goose would have ended the session with nothing built. - z-ai/glm-5.3-prime's Forge build declared its storage attributes as
sprintId: stringinstead of
sprintId: {type: string}. The scorer's deploy-readiness check raised AttributeError, so a paid run got no
verdict. - Forge's deploy check accepted entity indexes with more than one
rangeattribute. Forge allows only one
("This parameter can only have one attribute", storage-reference/entities-manifest). So the run pages of
qwen3.8-max-prime, fireworks/ember-1 and meta/muse-spark-1.3 said their apps would deploy when they would
not. The sprint-index checks already refused those indexes, so their scores are unchanged.
The fixes
- A streamed reply that ends at the output limit now always carries the truncation marker, even when it had no
text (reasoning only). - A reply cut at the output limit before any tool call is no longer treated as the final answer or as an empty
turn. goose tells the model the reply was cut off and asks it to continue in smaller steps, at most three
times in a row. The count resets after any tool call. This applies to every goose session, not only to
benchmarks. - The Forge scorer checks the shape of every manifest value it reads. A wrong shape fails the matching deploy
rule instead of crashing the scorer. Rule M7 now names attributes that are not{type: …}mappings, partition
or range values that are not lists, a range with more than one attribute, and indexed attributes of type
any.
What does not change
The benchmarks, contracts, fixtures and vendors. No Gauntlet or Forge score changes. A CLI replay of the
published Max Prime and Qwen3.8 Flash Forge builds gives 0.6953 and 0.7962 (published: 0.6954 and 0.7961). Only
the deploy-readiness text changes for builds with an invalid index. The Forge bench manifest is refrozen for
score_forge.py.
3.0.107 Reasoning details stay on the message, not inside the tool call
3.0.107 — reasoning details stay on the message, not inside the tool call
What happened
3.0.106 started sending OpenRouter reasoning details back with every assistant turn. They were also copied inside each
tool call of that turn. Most providers ignore the extra field; Fireworks rejects it, and fireworks/ember-1's two runs on
2026-10-06 ended on the second request with "400 Extra inputs are not permitted, field:
messages[2].tool_calls[0].reasoning_details".
The fix
Reasoning details now go back only at the message level, where OpenRouter reads them; the tool call keeps its other
fields (such as Gemini's thought signature) (goose d46425a).
What does not change
The benchmarks, contracts, fixtures, vendors and scorers (bench manifests unchanged).
3.0.106 OpenRouter reasoning details travel back with every tool call
3.0.106 — OpenRouter reasoning details travel back with every tool call
What happened
Gemini 3.8 Flash's Forge runs on 2026-10-06 ended on repeated empty answers (Google AI Studio after 47 messages,
Google Vertex later). Replaying the exact request showed an immediate empty "stop". The cause was in goose: OpenRouter
returns each turn's reasoning details (for Gemini 3, an encrypted thought signature) in a chunk before the tool call,
and goose's stream parser collected them but never attached them to the tool request, so the conversation was sent
back without any of them. Gemini 3 needs them across tool-calling turns.
The fix
A response's reasoning details now ride on its first tool request and go back with that assistant turn on every later
request (goose c99d3b0). This affects every OpenRouter model that returns reasoning details on tool calls.
What does not change
The benchmarks, contracts, fixtures, vendors and scorers (bench manifests unchanged).
3.0.105 OpenRouter's retry-after-in-flight 402 is waited out
3.0.105 — OpenRouter's "retry after in-flight requests settle" is waited out
What happened
meta/muse-spark-1.3's Forge run on 2026-10-06 ended at call 100 (about $12 spent) on an OpenRouter 402: "This request
would exceed your available credits given your current in-flight requests. Retry after in-flight requests settle, or
add credits." The account was topping itself up at the time. goose treated every 402 as fatal, the session ended, and a
session the provider ends is not scored.
The fix
A credits refusal whose own text says to retry (in-flight requests settling) now gets the same patient wait as a rate
limit (5 s doubling to 2 minutes, up to 30 minutes, the provider's Retry-After first). A plain "insufficient credits"
refusal still ends the session, and provider connection checks keep their quick retries (goose 34eaabe).
This release also carries 3.0.104's Gauntlet charge for a 3D scene whose heights are off the contract.
What does not change
The benchmarks, contracts, fixtures, vendors and scorers (bench manifests unchanged from 3.0.104).
3.0.104 A 3D scene whose heights are off the contract is scored, not refused
3.0.104 — a 3D scene whose heights are off the contract is scored, not refused
What happened
meta/muse-spark-1.3 finished its Gauntlet 7.2 build on 2026-10-06 at 68 calls and the scorer refused it: "stream
witness unavailable without a candidate cause". Its tower positions match the public layout exactly, but its heights
do not (its own debug digest: summed heights 21,764 against the contract's 25,028; 0 of 6 sampled heights right), so
none of the 11 clicks the probe made at the target's contract pixel brushed it. 3.0.103 charged a scene drawn off the
layout but not one whose heights are wrong, and it required the debug surface on every evaluation (Muse's was present
on 38 of 39).
The fix
The charge now covers a scene the app's own debug surface reports off the contract in layout OR in geometry over the
same record count, with the surface present on at least 90% of evaluations and the brush always readable. A digest
that differs over a different record count (a sync desync), a surface mostly absent, or a scene on the contract still
refuse (goose c26ab18).
What does not change
The benchmark task, contract, fixtures and vendor; Gauntlet 7.1 and Forge 1.0 scoring.
3.0.103 Broken 3D scenes and undeclared Forge resources are scored, not refused
3.0.103 — a 3D scene drawn off the public layout is scored, not refused
What happened
inclusionai/ling-3.0-flash finished its Gauntlet 7.2 build on 2026-10-06 and the in-app scorer refused it: "stream
witness unavailable without a candidate cause". Its vs7dbg debug surface was present on all 169 evaluations, but the
app draws its scene off the public layout (its own layout reports the first day as 2026-01-30 where the contract
says 2026-01-31, and its geometry digest is far off), so none of the 33 clicks the probe made at the target's
contract pixel brushed it. The owner's rule is that a broken build gets a bad score, never a refusal.
The fix
Gauntlet 7.2 now charges that case: the D1 decision corner and the stream pixel-witness rows score 0 with the
measured reason, when the app's own debug surface is present on every evaluation and reports a layout off the
contract and the probe's clicks at the target pixel went unbrushed. A correct or unmeasured layout, an unreadable
brush, a surface missed even once, or an early arm that reached the target still refuse (goose 8f24f08).
Also: Forge, a surface lost to the app's own undeclared resource
Ling 3.0 Flash's Forge app named resource widget in its dashboard widget without declaring it under resources
(real Forge rejects that manifest). The local UI host threw, all 35 UI-dependent rows were filed as harness failures
and the verdict could not be published. Those rows now score 0 with that reason when the app's own manifest proves the
cause; an error about a resource the manifest never references, or one it does declare, stays a harness failure
(goose b1a4a7e).
What does not change
The benchmark task, contract, fixtures and vendor; Gauntlet 7.1 and Forge 1.0 scoring. Builds that were refused on
this cause can be re-scored in the app with Retry scoring.
3.0.102 A pinned host's own context window is the run's limit
3.0.102 — a pinned OpenRouter host's own context window is the run's limit
What happened
Benchmark runs pinned to one OpenRouter host took their context window from OpenRouter's model listing. For
inclusionai/ling-3.0-flash the listing says 262,144 tokens, but the pinned DeepInfra bf16 host serves 131,072 and
reserves 32,768 of it for the answer. goose never compacted, and on 2026-10-06 the Gauntlet run ended at about call
21 when its prompt reached 102,198 tokens ("exceeds the model's maximum context length"), unscored.
The fix
When a run pins its host (OPENROUTER_PARAMETERS provider.order with allow_fallbacks false), the benchmark reads that
host's own endpoint listing and uses the smallest admitted window less the completion it reserves. The run's
model-limits.json records the pinned host and the usable window. A pin that names no host serving the model is
refused before any model call (goose 796a612).
What does not change
The benchmarks, contracts, fixtures, vendors and scorers. Unpinned runs keep the listing's window. The bench release
manifests are refrozen for run_build.py only.
3.0.101 Forge really can benchmark a model served on this Mac
3.0.101 — Forge really can benchmark a model served on this Mac
What happened
3.0.99 taught the Forge sandbox's provider list that an endpoint on this machine needs no relay, but the
sandbox's start-up check still proved "the provider answers through the relay" by taking the first relayed
host. With our own fine-tune served by the LeanZero MLX engine as the only provider, that list was empty and
the run stopped before its first model call (IndexError).
The fix
When every selected provider is served on this machine, the start-up check proves one of them answers
directly from inside the sandbox (loopback), alongside the two existing proofs that the internet is closed.
A local endpoint that does not answer refuses the run with that reason (goose, bench_isolation.py).
What does not change
The benchmarks, contracts, fixtures, vendors and scorers. Cloud providers keep the relay and its proof.
The bench release manifests are refrozen for bench_isolation.py and its fence test only.
3.0.100 A rate limit is waited out, not fatal
3.0.100 — a rate limit is waited out, not fatal
What happened
qwen/qwen3.8-flash is served on OpenRouter by one host (Alibaba). On 2026-10-05 it answered 429 "Rate limit
exceeded: Provider returned error" at call ~148 of a 150-call Gauntlet run and 68 seconds into the Forge run
that followed. goose resent each failed request three times over about seven seconds, then ended the session;
a session the provider ends is not scored, so neither run produced a result. Ling 3.1 Flash died the same way on
2026-10-03.
The fix
A rate limit now has its own budget: goose waits the provider's Retry-After when it names one, otherwise 5 s
doubling to 2 minutes, for up to 30 minutes in total (GOOSE_RATE_LIMIT_MAX_WAIT_SECS, 0 restores the old
behaviour), before the ordinary retries decide. Provider connection checks keep the quick retries, so Refresh
providers still answers in seconds (goose d0ff23d).
What does not change
The benchmarks, their contracts, fixtures, vendors and scorers are byte-identical; the bench release manifests
are unchanged. The 150-call budget is unchanged: a resent request is not a new model call.
3.0.99 Forge can benchmark a model served on this Mac
3.0.99 — Forge can benchmark a model served on this Mac
What happened
A single-model Forge run whose provider points at this machine (our own Qwen3.8-27B fine-tune served by the
LeanZero MLX engine at http://127.0.0.1:8096 through the OpenAI provider) was refused before the model was
called: the Forge sandbox's provider relay accepted HTTPS endpoints only, and a run whose only provider was
local counted as "selects no cloud provider". Gauntlet runs on an open network and was not affected.
The fix
The Forge sandbox already lets the entrant reach this machine directly (loopback outbound) and goose keeps
loopback traffic off the relay. A provider endpoint on 127.0.0.1 / localhost / ::1 now needs no relay entry and
counts as a reachable model. A plain-http endpoint on any other host is still refused (goose 30f0da1).
What does not change
The benchmark task, the contracts, the fixtures, the vendor and both scorers are byte-identical. The bench
release manifests are refrozen only for bench_isolation.py and its fence test.