Repository navigation
Releases: beamivalice/sushi
Releases · beamivalice/sushi
Release list
sushi v1.2.1
v1.2.1 — Safer prompt cache and server hardening
- Prompt cache: two sushi processes on the same model no longer share one SSD cache folder, which could return
another session's answer; a restart on a nearly full disk keeps the cache, and loading a second model no longer
deletes the first one's entries. - Crash fixes: an empty
/v1/completionsprompt, deeply nested JSON and oversized WebSocket messages are refused
with a 400 instead of stopping the server. - GLM-5.3-Flash: after a client disconnects, the next request on a streamed GLM no longer fails;
sushi pull
now fetches the DFlash2 assistant, and deleting its BF16 source keeps DFlash2 on. - Tools and structured output: earlier tool calls with empty or non-object arguments render in the model's own
format,anyOf/oneOf/$refparameters get their real types, and JSON-schema output stops cleanly on numbers and
bounded arrays;logit_bias: nullis accepted again. - CLI:
sushi launchquotes model names,sushi servewithout--modelhonours the sampling flags,update
no longer logs the API key, andsushi runaccepts long pasted lines and filters terminal escapes from model output. - mlx-serve: the guest manifest now lists GLM-5.3-Flash.
sushi v1.2.0
v1.2.0 — GLM-5.3-Flash, 32 GB streaming, zero-RAM prompt cache
- GLM-5.3-Flash: the new Sushi-2.4bpw pack runs on 128 GB Macs (KLD 0.074 against the BF16 model) with image and
video input, the full 1M context and DFlash2 drafting: 42–55 tok/s decode on an M5 Max. Up to four requests decode
together. - Streaming with MTP:
--ssd-budget-gbstreams any model's experts from the SSD, so Qwen3.8 Sushi-2bpw runs on a
32 GB Mac at about 20 tok/s. - Prompt cache with 0 GB of RAM: prompts are reused across turns from an SSD cache that sizes itself (up to 20 GB);
keeping them in RAM is now opt-in with--prefix-cache-mem. - Faster: MiMo decodes about 12% and prefills about 15% faster, Qwen3.8-Flash-Next prefills faster, and concurrent
requests decode together on every model. - Agents and API:
sushi launch grokis new and opencode 2.x works again; streamed and non-streamed answers match
byte for byte; penalties,repetition_penaltyandignore_eoswork; GLM history renders exactly as its template. - Flags:
--mtp-min-depth/--mtp-max-depthreplace--mtp-depth; new--no-mtp-lookupand--gpu-warm-secs;
--wired-margin-gibdefaults to 4 GiB. - Reliability: long GLM sessions, memory pressure and restarts are handled cleanly, and an idle streamed server no
longer uses CPU.
Thanks @cnsiva (request-budget defaults, repetition_penalty), @jasontitus (unique response IDs, the M1–M4 GLM fix,
portable GLM tests), @ShoichiTect (idle SSD-read workers) and @gomezvd (MTP on streamed Qwen).
sushi v1.1.1
v1.1.1 — Long MiMo prompts and agent sessions
- MiMo long prompts are admitted again: a resident MiMo server now sizes requests against the GPU limit you set
(iogpu.wired_limit_mb), not against what other apps happen to leave free, so a long agent session no longer gets
"requires ~N MB GPU memory" on a 768k server; the prompt cache gives its memory back to a request that needs it. --prefill-chunkis a maximum: a request that does not fit at your chunk steps down to a narrower one instead
of being refused. 2048 is the recommended value; wider chunks cost memory without prefilling faster.- Warm agent turns stop spiking memory: a turn that reuses the cached conversation grows its KV buffers during the
prefill, one layer at a time, instead of all at once on the first reply token, so long agent sessions stay admitted. - Claude Code on a local model:
sushi launch claudekeeps each turn on one streamed request, and a request whose
client disconnects now stops generating instead of running on for nobody. - Homebrew gets each release right away:
brew upgrade sushisees a new version as soon as it is published.
sushi v1.1.0
- MiMo-V2.6-Flash: sushi's second model.
MiMo-V2.6-Flash-Sushi-2.3bpwserves text and image input from one
resident pack with native MTP, up to its full 1M-token context on a 128 GB Mac. - Sushi packs stream from SSD:
--ssd-budget-gb Nkeeps the trunk resident and streams the routed experts from
SSD, so a Mac with less memory than the pack can serve it; replies are identical to a resident load. Every Sushi
Qwen pack and MiMo-V2.6-Flash-Sushi-2.3bpw stream, and Sushi-2bpw serves on a 32 GB M1 Max at a 20 GB budget. - Faster Flash-Next: on an M5 Max, Sushi-4bpw decodes 8% faster (83 -> 89 tok/s) and prefills a 10k-token prompt
19% faster (1,796 -> 2,139 tok/s). oMLX's tensor-unit sparse attention now serves prefill from the first sparse
chunk, GDN prefill runs a software-pipelined recurrence, batched decode overlaps GPU work with graph building, and
--prefill-decode-sharekeeps decoders moving while a long prompt prefills. Thanks @STRML and @cowboycoderhq. - Better and smaller packs: new expert weights for Sushi-4bpw (KLD 0.0632 -> 0.0592) and Sushi-3bpw
(0.1047 -> 0.1036), and Sushi-2bpw for 48 GB Macs. - Live sessions on
/metrics.json: every in-flight request and every cached conversation, with its phase,
context against the model's limit and the GPU memory its KV holds. Thanks @yoyo930021 and @ddalcu. - Updates itself:
sushi updateinstalls the newest release after checking its SHA-256 and signature and keeps
the old one for--rollback; a daily check, one-click update from the chat page, orbrew install beamivalice/tap/sushi. - Fixes: MiMo long-context memory returns to the OS, SSD prompt-cache restores keep one copy, JSON-constrained
logprobs pair with their tokens, and a failed model load can be retried after a rescan. Thanks @brandondyal and
@jasontitus; the EXL3 engine is now a module mlx-serve builds against, thanks @ddalcu.
sushi v1.0.5
v1.0.5 — Sushi-2.6bpw at full speed
- Sushi-2.6bpw for 64 GB Macs, as fast as Sushi-3bpw: the new Qwen3.8-Flash-Next pack carries 43.95 GiB of
weights, 5.4 GiB less than Sushi-3bpw, so a 64 GB Mac serves 250k tokens of context at 8-bit KV. Its experts now run
on the same fast kernels as Sushi-3bpw's: 18% faster decode and 12% faster prompts on an M5 Max, output unchanged. - Faster Flash-Next decoding: when a reply copies text already in the conversation (a file returned with an edit,
a tool call carrying a file), speculative decoding drafts from that text, 16-21% faster file edits and 9-11% faster
file writes; linear attention, hyper-connections and the sparse-attention indexer also take fewer GPU dispatches.
Output is unchanged; several of these are ported from mlx-serve, thanks @STRML. - Long conversations stay cached: by default the RAM prompt cache holds a whole conversation where memory allows,
a session longer than the cache keeps the longest prefix that fits, and sessions past about 250k tokens restore
from the SSD cache with half the memory. Ported in part from mlx-serve, thanks @STRML, @brandondyal and
@celestial-rose. - Forced tool calls that work:
tool_choicerequired(Anthropicany) or a named function now makes
Qwen3.8-Flash-Next call one of the declared tools on every API, after its thinking; naming an undeclared function
is a 400. Streamed and non-streamed thinking are now the same text, and a continued reply resumes after its
closed think block. --preserve-thinking on|off: keep every turn's thinking in the prompt (the default) or only the latest turn's,
per model in model-settings.json or per request withchat_template_kwargs.preserve_thinking.sushi run qwen3.8-flash-next:sushi runandsushi pullname the Sushi packs (:2.6bpw,:4bpw, 3bpw by
default) instead of models sushi does not serve.- Thanks, @jasontitus: for this release's faster MTP verification on Flash-Next (hyper-connection weights read
once per group of draft tokens), the fix that stops long prompts being refused while the GPU finishes earlier work,
and the build-from-source docs, and, belatedly, for v1.0.4's 4x faster prompts on M1–M4 Macs and parallel n-gram
reads on 64 GB Macs.
sushi v1.0.4
v1.0.4 — Faster prompts and browser chat
- Faster prompts on M1–M4: Macs without the M5's neural accelerators read prompts about 4x faster on the EXL3
packs (M2 Max, 3–4k-token prompts, default settings). - Faster Flash-Next prompts on 64 GB Macs: when the n-gram table cannot stay in memory beside the model, prompt
processing reads it in parallel by default instead of one row at a time. - Smaller contexts need less free memory to load: resident Flash-Next EXL3 packs without separate sidecars or
ANE now size load headroom from the chosen context instead of always asking for 7 GB above the weights. - Chat in your browser:
sushi serveandsushi runserve a chat page athttp://127.0.0.1:12345/that streams
replies, shows the model's thinking, takes images for vision models and keeps your conversations in the browser. /cd <folder>insushi run: moves the folder the file tools and relative/imagepaths read from, the prompt
always shows that folder and whether tools are on, and a model that asks for a file outside it now suggests/cd.
sushi v1.0.3
v1.0.3 — Hotfixes
- Prompt cache: re-packing a model in place no longer restores stale SSD cache entries, and the RAM cache stays
within its cap when every entry is in use. - Video and labels: multi-part videos that fit are no longer refused, and
/v1/modelsnames an EXL3 pack's
expert rate beside its dense width.
sushi v1.0.2
v1.0.2 — Hotfixes
- Stability: two cached conversations can no longer share a prompt-cache key, a failed long-context cache copy no
longer frees memory twice, and very low temperatures behave the same with and without MTP. - Memory and loading: two resident models no longer over-commit memory at admission, and packs with invalid EXL3
rate stamps are refused by name.
sushi v1.0.1
v1.0.1 — Hotfixes
- Tool calls and streaming: streamed tool calls are always valid JSON, the reasoning budget applies to tool replies
and Anthropic streams, and a disconnected Responses request is no longer stored as completed. - Prefix cache and loading: fixes for SSD cache restores and failed cache writes, and malformed pack configs are
refused by name.
sushi v1.0.0
v1.0.0 — Qwen3.8-Flash-Next, sushi-packed
Add support for 2 sushi models
- Qwen3.8-Flash-Next-Sushi-3bpw (requires 64 GB+)
- Qwen3.8-Flash-Next-Sushi-4bpw (requires 96 GB+)
Quality
KLD against the bf16 model: 16 prompts x 512 tokens scored to the first EOS, kv8, every pack run by the same sushi
build. The light rings are the sushi packs with a 4-bit n-gram table (Sushi-3bpw ships that table; either table works
with either pack). Numbers: docs/quality-kld.md.
Speed
Sushi-3bpw on an M5 Max 128 GB, sushi v1.0.0 release candidate (build 725b76c): --ctx-size 1048576 --kv-quant 8 --mtp, llmprobe --bench-only, quiet box.

