Repository navigation
Releases: NathanL15/Hex-CLI
Release list
v2.22.5
A patch release: nothing the model reads changes.
- The start-up prime no longer forces a rebuild of a cache the first turn
can extend. A server that already served a session holds this prompt from
that session's end-of-turn prewarm, and the forced prime rebuilt it anyway.
It now asks without forcing first, and forces only an empty cache or one
longer than the prompt plus the 600-token Rewind discard limit. A new
session's first turn on an already-used server, question typed 5 s after
start: 13.7–14.7 s on 2.22.3, 6.0–6.3 s on 2.22.4, 1.9–2.3 s now (4
runs each). On a freshly started server the first two turns are unchanged
at 11.7–12.3 s in total. After a Hex upgrade that changes the prompt, the
first turn of the first session pays one rebuild.
v2.22.4
A patch release: nothing the model reads changes.
- The start-up prime warms the prompt the first turn actually sends. It ran
before any turn had set the session config, so it built the default
prompt (date and working directory in its second line), which shared 88
characters with the first turn's stable prompt. The prime was discarded
and the first turn rebuilt the dialog. Measured on a server that had
already served a session, with the question typed 5 s after the prime
(the owner's median over 85 REPL sessions): first model call 13.7–14.7 s
with one Rewind failure every time, now 6.0–6.3 s with none (4 paired
runs). On a freshly started server the first Rewind fails once whatever
is cached (a runtime quirk, measured the same night), so the first turn
after a server start still pays one rebuild.
v2.22.3
A patch release: nothing the model reads changes.
- A reused keep-alive connection now takes each request's own timeout. It
kept the timeout of the request that opened it, so a non-streamed model
call made after the 5 s end-of-turn prewarm or the 3 s backend ping was
cut off at 5 or 3 s and reported as "backend unreachable: timed out". Of
the 237 recorded eval runs marked invalid for a timeout, 206 ended at 3 or
5 s and 2 at the real 300 s limit, so most of the "platform" invalid runs
on record were this. The interactive REPL streams replies over another
path and was not affected; the eval runner anduse_streaming: falsewere.
v2.22.2
A patch release: nothing the model reads changes.
- Each eval trace records the server-log counts for its own run (requests,
Rewind failures, busy-slot retries), and each model call its
finish_reason. Until now those were only reported per arm. - The two exact tests are checked against SciPy 1.16's values. They matched
to 1e-15 on every 2×2 table up to 12 runs per arm and every discordant
pair up to 40 each; 21 arm-sized points are pinned in the tests. - Measured, nothing changed: Hex's token estimate against the real Qwen
tokenizer over 134 reconstructed requests from the chat log runs a median
3.5 % high (−2.7 % to +7.5 %), and the server's 3.6–5.4 % high, so both
err on the safe side and exact counting would buy about 3 % of headroom.
v2.22.1
A patch release: only the eval harness changes.
python evals/cases_smoke.pyruns ten cases from the pinned gate set
(SMOKE_SET) instead of its own ten. The old set heldfactual-1(19/24),
agentic-2(18/24),agentic-3(18/20) andlint-1(20/21), so
unchanged code scored 10/10 about half the time and most releases
re-measured one of them. Each of the new ten passed every run in the two
arms that pinned the gate set and in the 16 arms since (44–51 runs each).
The new set went 10/10 on its first live run. The extended suite is
unchanged.
v2.22.0
A minor release: a tool result the model reads changes (edit_file when
new_string carries a literal \n and the file would not parse).
Gate: PASS on the pinned 24-case set after one recheck:
agentic-1
had all five runs invalid ("backend unreachable", the server stalled when
free memory hit 1.4 GB), and re-run on a fresh server it was 6/6. Own arm,
seed 20261011, against 2.21.0's arm: run-level 184/234 vs 186/233,
−1.2 %, Fisher p = 0.82; pass^k 34 → 32 of 49, McNemar p = 0.73; 11
invalid runs of 245. The decode never fired in the arm: none of its 48
edit_filecalls had a literal\ninnew_string, so every edit result
was byte-identical to 2.21.0 and the arm-level movement is run-to-run
noise. The gate shows nothing broke; the benefit is not yet measured live.
- When an edit to a
.pyor.jsonfile would leave it unparseable and
new_stringhas a literal\nbut no real line break,edit_filetries
the decodednew_string(the full decode, then for Python one that keeps a
quote-adjacent escape) and keeps it only if the file then parses. 139 of
1,746 recordededit_filecalls put a literal\ninnew_string, 135
with no real line break, mostly inself-correct-1,agentic-3and
regression-anchor-1, and the 2.21.0 arm lost a run of each of the latter
two to it. An escape that parses as written is never touched, and other
file types are left alone.
v2.21.0
A minor release: a tool result the model reads changes (edit_file when
old_string is not in the file).
Gate: PASS on the pinned 24-case set after one recheck:
error-recovery-3missed one of five runs, and that run never called
edit_file(it reported a "permission error" afterfind_files); re-run on
a fresh server it was 6/6. Own arm, seed 20261010, against 2.20.0's arm:
run-level 189/237 vs 192/245, +1.4 %, Fisher p = 0.74; pass^k 32 → 34 of
49, McNemar p = 0.73; 8 invalid runs of 245 (backend timeouts, 2.2 GB
free).claims-2, one of the three cases behind most of the recorded
misses, went 2/5 → 5/5.self-correct-1lost one run to the known
double-escapednew_string(a literal\nlanded on the first edit,
before any miss), andregression-anchor-1lost one to the same defect
with no miss at all. Smoke 9/9 valid on a fresh server.
- A missed
edit_fileon a file of up to 1,500 characters returns the
file's whole current text instead of the closest region. Of 1,279 recorded
edit_filecalls, 786 were made without reading the file that turn and
failed 45 % of the time against 9 % after a read, and 115 of the 291
misses came straight after another miss on the same file: the retry was
rebuilt from memory again. Every fixture behind those misses is under 200
characters. In the arm the message fired 16 times, intests-claim-1,
self-correct-1,claims-2andagentic-3, and the next edit landed
after 7 of them (44 %), against 14 of 39 (36 %) in the three arms before
it; at that count the recovery rate is not yet a measured effect. A larger
file keeps the closest-region report.
v2.20.1
A patch release: nothing the model reads changes. The turn's return value
and the history entry for a stopped turn are byte-identical to 2.20.0.
- A turn the loop stops now says which tool failed, on which file, and
whether any file changed: "Stopped after edit_file failed three times on
calculator.py; no file was changed." instead of "Stopped: edit_file failed
three times in a row." 33 recorded turns ended on a loop or step-limit stop
with only the rawold_string was not founderror behind them, and none of
them said whether the file had been touched. The file clause comes from the
writes that landed this turn, not from the attempts, so a turn that wrote
notes.txtand then got stuck oncalculator.pynamesnotes.txt; a turn
that never tried to change a file gets no clause. - Off a terminal (
hex "..." | clip, scripts) a stopped turn printed
nothing at all. It prints the stop sentence. The raw tool output still goes
to the session history, where the next turn reads it.
v2.20.0
A minor release: a tool result the model reads changes (run_code on a
program that waits for the keyboard).
Gate: PASS on the pinned 24-case set: 21 cases held 5/5 in the arm;
casual-4,factual-4andmemory-1lost four, two and four runs to
backend timeouts and were re-run on a fresh server at six runs, 6/6 each.
Own arm, seed 20260918. Platform: 13 invalid of 245 runs (5.3 %), the
server's working set again above 6 GB with 1.6 GB free, so the ceiling
comparison carries the[PLATFORM]flag: run-level 402/514 vs 166/217,
−1.7 %, Fisher p = 0.63; pass^k 26 → 30 of 46 shared cases, McNemar
p = 0.13. The change did what it is for:run_codetimed out 0 times in
the arm against 10 in the 2.12.0 arm (all of themmake-py-1's
calculator waiting oninput()), the note fired twice, andmake-py-1
stayed 5/5. Smoke 10/10 on a fresh server.
run_codegives the program no keyboard, and says so when the program
asked for one. The child inherited the terminal's standard input, so a
program that callsinput()sat waiting on the owner's console until the
tool's timeout: on 2026-09-18 "create a python cli hilo guessing game and
run it" wrote a working game and then waited the full ten seconds, twice,
to report "TIMEOUT". Standard input is now closed for the child, so
input()raisesEOFErrorat once, and when it does the result adds one
line: the program reads from the keyboard,run_codehas none to give
it, and it has to be started in a terminal, with the command. On record:
12run_codetimeouts whose output ends in a prompt (ten of them the
make-py-1calculator) and 19 runs that already hitEOFErrorunder the
eval runner, which has no terminal; 0.6 % of recorded runs, all of them
this shape.
v2.19.0
A minor release: nudge text the model reads is new (the claimed-fix rule)
and two existing rules fire where they did not.
Gate: PASS on the pinned 24-case set: 21 cases held 5/5 in the arm;
missing-file-1,missing-file-2andregression-count-1lost one, three
and one runs to backend timeouts and were re-run on a fresh server at six
runs, 6/6 each. Own arm, seed 20260918. Platform: 16 invalid of 245 runs
(6.5 %) -- the model server's working set grew from 3 GB to 6.7 GB over
the hour and the machine fell to 1.2 GB free, so the ceiling comparison
carries the[PLATFORM]flag: run-level 391/502 vs 161/214, −2.6 %, Fisher p = 0.44;
pass^k 26 → 30 of 46 shared cases, McNemar p = 0.13; every ceiling drop
(livestate-11/4,tests-claim-10/2,claims-11/5) is on a case that
lost runs to timeouts. The new and widened rules did not fire once in
the arm, so the arm is evidence of no regression on the gated set and
the effect rests on the replay below. Smoke 10/10 on a fresh server.
- The intent nudge counts only a mutation that landed, catches a claimed
fix with no change behind it, and reads "cannot be created" however the
sentence goes on. Three gaps, all from the owner's 2026-09-18 session.
The loop had been passing "a tool with a path was called" as mutated,
which anedit_fileon a missing file satisfies, so when the turn ended
"A calculator app cannot be created without a matching file" right after
exactly that error, the create rule stayed silent -- and its wording only
knew "cannot be created in this environment" anyway. Later, "fix it" was
answered with "The file has been fixed by removing the malformed line"
after arun_codethat showed the SyntaxError and no edit at all; nothing
looked at that. Now mutated means a write, edit or append that did not
error; the denial wording matchescannot be (built|created|done|made)
with any continuation; and a request to fix, correct, update or change
something (an instruction, not a question) answered with "has been fixed",
"was corrected", "is now updated" and no mutation gets one nudge to make
the change or say plainly that nothing was changed. Replayed over the 206
real turns and 5,389 recorded runs with the loop's own semantics: on real
turns the nudge fires 11 times, all genuine (two new: the refusal above
and the claimed fix); on recorded runs the new rule fires on seven, five
of themambiguous-2answering "Update the file." with "Updated the file
as requested" and no tool call, two of themtests-claim-1claiming a
correction it never made. The question form is what keepsmemory-1
("which file did you fix?") out: its honest recall fired 14 times in the
first draft and zero now.