Skip to content

v0.0.302.14

Choose a tag to compare

@github-actions github-actions released this 02 Oct 03:30
· 17 commits to main since this release
ff44804

Codegraff v0.0.302.14

MiMo now answers without thinking unless you ask it to, MCP servers connect
in the background, sub-agents work the way they do on the codex route, and
shell waits and todo_write replies stop costing turns. Measured against the
previous release and other harnesses on the same models:

Evals

MiMo v2.6 Pro: 2.9x as fast as v0.0.302.13

The published v0.0.302.13 binary, this release and OpenCode 2 on the same
model, endpoint and key, interleaved over 11 coding and 10 MCP tasks, 2 runs
each:

Pass Wall per task Coding tasks MCP tasks Output tokens per task
v0.0.302.13 42/42 53.1s 15.8s 94.2s 2,353
v0.0.302.14 39/42 18.1s 11.5s 25.5s 544
OpenCode 2 42/42 30.9s 17.4s 45.7s 1,004

MiMo used to think at length before every action, and that was most of a
task's time. Without thinking it missed 3 of 42 runs: one reply was a JSON
fragment instead of a tool call, and one task twice reported the issues'
display numbers instead of their ids. /effort high turns thinking back on
for work that needs careful reading.

MiMo v2.6 Flash, thinking on against this release's default: 35.1s against
16.7s per task (OpenCode 2: 26.6s), 42/42 against 40/42.
Full results.

DeepSeek V4 Pro: graff against deepseek-harness

Both harnesses on DeepSeek's own API with one key, interleaved, 2 runs each:

Pass Wall per task Coding tasks MCP tasks Model calls per MCP task
graff 42/42 22.6s 6.4s 40.4s 5.0
deepseek-harness 0.2.0-rc.2 42/42 27.7s 7.4s 50.2s 9.2

graff takes 19% less time over the 21 tasks. On the sub-agent suite it passed
9 of 10 runs in 63.7s per task against 8 of 10 in 70.4s, and doing the same
work without delegation 27.7s against 38.1s.
Full results.

gpt-6-astra: graff against the Codex app server, with sub-agents

On the same model and ChatGPT account, graff is faster on every suite: 20% on
the coding tasks, 30% on the MCP tasks, 18% when the tasks ask for sub-agents
and 6% when they do not, at 16% to 73% lower cost on the app server's own
reported usage.
Full results.

MiMo thinks only at high effort

graff's default effort is medium, and it used to turn MiMo's thinking on
for every request. minimal, low and medium now send thinking Off on the
Chat and Responses wires; high and above turn it on, and so do a worker's
effort pin and Jev's high. The picker, status line and ACP show the default
as Off (ADR 0235, #1453).

MCP connects in the background

Every session but --json connects MCP servers in the background, so a slow
server no longer holds the first message. Each request merges the servers
that have finished, /mcp names the ones still connecting, and ACP
session/new waits at most 250 ms for a client's servers
(ADR 0230, #1445).

Sub-agents at parity with the codex route

  • write_file creates a missing directory instead of refusing and making
    the model regenerate the file.
  • agent_output takes ids, so collecting N children is one call.
  • A child starts with the user's prompts from its parent's history, runs at
    its parent's live effort, and can ask Jev for its own.
  • Read-only children on macOS run any command under a sandbox that refuses
    writes in the project and network access, so they can compute instead of
    reporting partial results.
  • Each finished child writes a subagent trace line, and graff-evals has a
    subagents suite with its subagents-solo counterfactual
    (ADR 0232, #1445).

Fewer wasted calls

  • Interactive sessions ignore a shorter shell timeout, so a quick command is
    no longer parked after a second and the turn no longer ends. todo_write
    replies with counts instead of repeating the list
    (ADR 0233, #1451).
  • The shell tool names the OS, so models write BSD sed -i '' on macOS. The
    codedb guard lets in-place sed edits and small whole-file reads through;
    searches and big reads still go to codedb
    (ADR 0234, #1452).
  • rlm takes JSON literals and binds them, each() maps plain values, a
    lone object argument is the keyword arguments, and a refused statement lists
    the forms that work
    (ADR 0236, #1454).