Skip to content

Releases: NathanL15/Hex-CLI

v2.22.5

Choose a tag to compare

@NathanL15 NathanL15 released this 11 Oct 07:20

A patch release: nothing the model reads changes.

  • The start-up prime no longer forces a rebuild of a cache the first turn
    can extend. A server that already served a session holds this prompt from
    that session's end-of-turn prewarm, and the forced prime rebuilt it anyway.
    It now asks without forcing first, and forces only an empty cache or one
    longer than the prompt plus the 600-token Rewind discard limit. A new
    session's first turn on an already-used server, question typed 5 s after
    start: 13.7–14.7 s on 2.22.3, 6.0–6.3 s on 2.22.4, 1.9–2.3 s now (4
    runs each). On a freshly started server the first two turns are unchanged
    at 11.7–12.3 s in total. After a Hex upgrade that changes the prompt, the
    first turn of the first session pays one rebuild.

v2.22.4

Choose a tag to compare

@NathanL15 NathanL15 released this 11 Oct 06:50

A patch release: nothing the model reads changes.

  • The start-up prime warms the prompt the first turn actually sends. It ran
    before any turn had set the session config, so it built the default
    prompt (date and working directory in its second line), which shared 88
    characters with the first turn's stable prompt. The prime was discarded
    and the first turn rebuilt the dialog. Measured on a server that had
    already served a session, with the question typed 5 s after the prime
    (the owner's median over 85 REPL sessions): first model call 13.7–14.7 s
    with one Rewind failure every time, now 6.0–6.3 s with none (4 paired
    runs). On a freshly started server the first Rewind fails once whatever
    is cached (a runtime quirk, measured the same night), so the first turn
    after a server start still pays one rebuild.

v2.22.3

Choose a tag to compare

@NathanL15 NathanL15 released this 11 Oct 06:27

A patch release: nothing the model reads changes.

  • A reused keep-alive connection now takes each request's own timeout. It
    kept the timeout of the request that opened it, so a non-streamed model
    call made after the 5 s end-of-turn prewarm or the 3 s backend ping was
    cut off at 5 or 3 s and reported as "backend unreachable: timed out". Of
    the 237 recorded eval runs marked invalid for a timeout, 206 ended at 3 or
    5 s and 2 at the real 300 s limit, so most of the "platform" invalid runs
    on record were this. The interactive REPL streams replies over another
    path and was not affected; the eval runner and use_streaming: false were.

v2.22.2

Choose a tag to compare

@NathanL15 NathanL15 released this 11 Oct 00:38

A patch release: nothing the model reads changes.

  • Each eval trace records the server-log counts for its own run (requests,
    Rewind failures, busy-slot retries), and each model call its
    finish_reason. Until now those were only reported per arm.
  • The two exact tests are checked against SciPy 1.16's values. They matched
    to 1e-15 on every 2×2 table up to 12 runs per arm and every discordant
    pair up to 40 each; 21 arm-sized points are pinned in the tests.
  • Measured, nothing changed: Hex's token estimate against the real Qwen
    tokenizer over 134 reconstructed requests from the chat log runs a median
    3.5 % high (−2.7 % to +7.5 %), and the server's 3.6–5.4 % high, so both
    err on the safe side and exact counting would buy about 3 % of headroom.

v2.22.1

Choose a tag to compare

@NathanL15 NathanL15 released this 10 Oct 23:17

A patch release: only the eval harness changes.

  • python evals/cases_smoke.py runs ten cases from the pinned gate set
    (SMOKE_SET) instead of its own ten. The old set held factual-1 (19/24),
    agentic-2 (18/24), agentic-3 (18/20) and lint-1 (20/21), so
    unchanged code scored 10/10 about half the time and most releases
    re-measured one of them. Each of the new ten passed every run in the two
    arms that pinned the gate set and in the 16 arms since (44–51 runs each).
    The new set went 10/10 on its first live run. The extended suite is
    unchanged.

v2.22.0

Choose a tag to compare

@NathanL15 NathanL15 released this 10 Oct 23:09

A minor release: a tool result the model reads changes (edit_file when
new_string carries a literal \n and the file would not parse).

Gate: PASS on the pinned 24-case set after one recheck: agentic-1
had all five runs invalid ("backend unreachable", the server stalled when
free memory hit 1.4 GB), and re-run on a fresh server it was 6/6. Own arm,
seed 20261011, against 2.21.0's arm: run-level 184/234 vs 186/233,
−1.2 %, Fisher p = 0.82; pass^k 34 → 32 of 49, McNemar p = 0.73; 11
invalid runs of 245. The decode never fired in the arm: none of its 48
edit_file calls had a literal \n in new_string, so every edit result
was byte-identical to 2.21.0 and the arm-level movement is run-to-run
noise. The gate shows nothing broke; the benefit is not yet measured live.

  • When an edit to a .py or .json file would leave it unparseable and
    new_string has a literal \n but no real line break, edit_file tries
    the decoded new_string (the full decode, then for Python one that keeps a
    quote-adjacent escape) and keeps it only if the file then parses. 139 of
    1,746 recorded edit_file calls put a literal \n in new_string, 135
    with no real line break, mostly in self-correct-1, agentic-3 and
    regression-anchor-1, and the 2.21.0 arm lost a run of each of the latter
    two to it. An escape that parses as written is never touched, and other
    file types are left alone.

v2.21.0

Choose a tag to compare

@NathanL15 NathanL15 released this 10 Oct 06:54

A minor release: a tool result the model reads changes (edit_file when
old_string is not in the file).

Gate: PASS on the pinned 24-case set after one recheck:
error-recovery-3 missed one of five runs, and that run never called
edit_file (it reported a "permission error" after find_files); re-run on
a fresh server it was 6/6. Own arm, seed 20261010, against 2.20.0's arm:
run-level 189/237 vs 192/245, +1.4 %, Fisher p = 0.74; pass^k 32 → 34 of
49, McNemar p = 0.73; 8 invalid runs of 245 (backend timeouts, 2.2 GB
free). claims-2, one of the three cases behind most of the recorded
misses, went 2/5 → 5/5. self-correct-1 lost one run to the known
double-escaped new_string (a literal \n landed on the first edit,
before any miss), and regression-anchor-1 lost one to the same defect
with no miss at all. Smoke 9/9 valid on a fresh server.

  • A missed edit_file on a file of up to 1,500 characters returns the
    file's whole current text instead of the closest region. Of 1,279 recorded
    edit_file calls, 786 were made without reading the file that turn and
    failed 45 % of the time against 9 % after a read, and 115 of the 291
    misses came straight after another miss on the same file: the retry was
    rebuilt from memory again. Every fixture behind those misses is under 200
    characters. In the arm the message fired 16 times, in tests-claim-1,
    self-correct-1, claims-2 and agentic-3, and the next edit landed
    after 7 of them (44 %), against 14 of 39 (36 %) in the three arms before
    it; at that count the recovery rate is not yet a measured effect. A larger
    file keeps the closest-region report.

v2.20.1

Choose a tag to compare

@NathanL15 NathanL15 released this 10 Oct 05:24

A patch release: nothing the model reads changes. The turn's return value
and the history entry for a stopped turn are byte-identical to 2.20.0.

  • A turn the loop stops now says which tool failed, on which file, and
    whether any file changed: "Stopped after edit_file failed three times on
    calculator.py; no file was changed." instead of "Stopped: edit_file failed
    three times in a row." 33 recorded turns ended on a loop or step-limit stop
    with only the raw old_string was not found error behind them, and none of
    them said whether the file had been touched. The file clause comes from the
    writes that landed this turn, not from the attempts, so a turn that wrote
    notes.txt and then got stuck on calculator.py names notes.txt; a turn
    that never tried to change a file gets no clause.
  • Off a terminal (hex "..." | clip, scripts) a stopped turn printed
    nothing at all. It prints the stop sentence. The raw tool output still goes
    to the session history, where the next turn reads it.

v2.20.0

Choose a tag to compare

@NathanL15 NathanL15 released this 18 Sep 11:34

A minor release: a tool result the model reads changes (run_code on a
program that waits for the keyboard).

Gate: PASS on the pinned 24-case set: 21 cases held 5/5 in the arm;
casual-4, factual-4 and memory-1 lost four, two and four runs to
backend timeouts and were re-run on a fresh server at six runs, 6/6 each.
Own arm, seed 20260918. Platform: 13 invalid of 245 runs (5.3 %), the
server's working set again above 6 GB with 1.6 GB free, so the ceiling
comparison carries the [PLATFORM] flag: run-level 402/514 vs 166/217,
−1.7 %, Fisher p = 0.63; pass^k 26 → 30 of 46 shared cases, McNemar
p = 0.13. The change did what it is for: run_code timed out 0 times in
the arm against 10 in the 2.12.0 arm (all of them make-py-1's
calculator waiting on input()), the note fired twice, and make-py-1
stayed 5/5. Smoke 10/10 on a fresh server.

  • run_code gives the program no keyboard, and says so when the program
    asked for one. The child inherited the terminal's standard input, so a
    program that calls input() sat waiting on the owner's console until the
    tool's timeout: on 2026-09-18 "create a python cli hilo guessing game and
    run it" wrote a working game and then waited the full ten seconds, twice,
    to report "TIMEOUT". Standard input is now closed for the child, so
    input() raises EOFError at once, and when it does the result adds one
    line: the program reads from the keyboard, run_code has none to give
    it, and it has to be started in a terminal, with the command. On record:
    12 run_code timeouts whose output ends in a prompt (ten of them the
    make-py-1 calculator) and 19 runs that already hit EOFError under the
    eval runner, which has no terminal; 0.6 % of recorded runs, all of them
    this shape.

v2.19.0

Choose a tag to compare

@NathanL15 NathanL15 released this 18 Sep 10:28

A minor release: nudge text the model reads is new (the claimed-fix rule)
and two existing rules fire where they did not.

Gate: PASS on the pinned 24-case set: 21 cases held 5/5 in the arm;
missing-file-1, missing-file-2 and regression-count-1 lost one, three
and one runs to backend timeouts and were re-run on a fresh server at six
runs, 6/6 each. Own arm, seed 20260918. Platform: 16 invalid of 245 runs
(6.5 %)
-- the model server's working set grew from 3 GB to 6.7 GB over
the hour and the machine fell to 1.2 GB free, so the ceiling comparison
carries the [PLATFORM] flag: run-level 391/502 vs 161/214, −2.6 %, Fisher p = 0.44;
pass^k 26 → 30 of 46 shared cases, McNemar p = 0.13; every ceiling drop
(livestate-1 1/4, tests-claim-1 0/2, claims-1 1/5) is on a case that
lost runs to timeouts. The new and widened rules did not fire once in
the arm
, so the arm is evidence of no regression on the gated set and
the effect rests on the replay below. Smoke 10/10 on a fresh server.

  • The intent nudge counts only a mutation that landed, catches a claimed
    fix with no change behind it, and reads "cannot be created" however the
    sentence goes on. Three gaps, all from the owner's 2026-09-18 session.
    The loop had been passing "a tool with a path was called" as mutated,
    which an edit_file on a missing file satisfies, so when the turn ended
    "A calculator app cannot be created without a matching file" right after
    exactly that error, the create rule stayed silent -- and its wording only
    knew "cannot be created in this environment" anyway. Later, "fix it" was
    answered with "The file has been fixed by removing the malformed line"
    after a run_code that showed the SyntaxError and no edit at all; nothing
    looked at that. Now mutated means a write, edit or append that did not
    error; the denial wording matches cannot be (built|created|done|made)
    with any continuation; and a request to fix, correct, update or change
    something (an instruction, not a question) answered with "has been fixed",
    "was corrected", "is now updated" and no mutation gets one nudge to make
    the change or say plainly that nothing was changed. Replayed over the 206
    real turns and 5,389 recorded runs with the loop's own semantics: on real
    turns the nudge fires 11 times, all genuine (two new: the refusal above
    and the claimed fix); on recorded runs the new rule fires on seven, five
    of them ambiguous-2 answering "Update the file." with "Updated the file
    as requested" and no tool call, two of them tests-claim-1 claiming a
    correction it never made. The question form is what keeps memory-1
    ("which file did you fix?") out: its honest recall fired 14 times in the
    first draft and zero now.