Skip to content

MTPLX 2.12.1

Choose a tag to compare

@youssofal youssofal released this 02 Oct 22:18
· 7 commits to main since this release

MTPLX 2.12.1

MTPLX 2.12.1 is a stability release for long coding sessions. It rewrites the memory guard that decides whether a request may run: every request is now priced before its prompt is read, idle conversations are released before anything is refused, and a refusal says what holds the memory and when a retry can work. On 128 GB Macs the engine now leaves more room for your other apps, a Pi compaction no longer waits behind the idle conversation it summarizes, and Flash-Next keeps its saved conversation through screenshots and its compiled verifier past 140K tokens.

On October 2, on an M5 Max with 128 GB, a Pi session on Flash-Next Optimized Speed grew to 204,981 tokens with 12 screenshots, with no errors and no memory stops. 86 of its 87 turns restored from RAM, and all 50,769 verify steps ran compiled. The one turn that read the conversation again hit a defect in that build, which is fixed in this release. A real Pi compaction of that session was served with no refusal and with swap flat. The release also fixes the leak behind "Failed to create Metal shared event" (#544), JSON schema output that looped on whitespace (#547), NaN output with 4-bit and 8-bit KV cache (#526), an app badge stuck on Degraded (#528) and mtplx start downloading a model you had already picked (#573). Every real-model check ran on that one Mac, and Known issues lists what is untested or still open.

Highlights

What you do 2.12.0 2.12.1
Let Pi compact a long Flash-Next session on a 128 GB Mac (#567) The summary request was refused with a 507 13 times in a row while the idle conversation kept its memory, until a restart Pi's two summary requests (138,774 and 21,597 tokens) were served with no refusal, and swap stayed flat (October 2)
Keep working right after that compaction In one report, a memory abort repeated on every resume until a restart The next turn (33,608 tokens and a screenshot) was served, and the turn after it reused 33,752 tokens from RAM (October 2)
Work in Pi with screenshots in the conversation A turn that changed anything after a screenshot resumed at best before the first screenshot, and could read the whole conversation again 86 of 87 turns restored from RAM and 99.2% of all prompt tokens were reused, with 12 screenshots, up to 204,981 tokens (October 2)
Send a screenshot after a tool call from a streaming client, as OpenCode does (#581) 3 of 5 streaming requests failed with "vision requests must not use the session bank without content-keyed ids" 5 of 5 were served (October 2)
Keep a Flash-Next conversation going past 140K tokens The compiled verifier first copied the attention cache into a padded buffer, 4.2 GB at 147,396 tokens, and verified eagerly when memory could not fund the copy The cache grows in place, one layer at a time (0.36 to 0.57 GB at that size), and 50,769 of 50,769 verify steps ran compiled (October 2)
Run long non-streaming jobs on Flash-Next (#544) About 2,900 Metal events were held during one 4,096-token answer, and the server failed after 141K to 263K generated tokens 1 to 9 events through three 4,096-token answers (October 2). A run past 140K generated tokens has not been repeated
Run at a 90 GiB memory limit on a 128 GB Mac Memory outside the GPU allocator was forgiven up to 22 GiB, so the process could still reach 112 GiB Forgiven up to 8 GiB, so the guard holds the process to 98 GiB (October 2)
Ask Flash-Next for strict JSON (#547) 7 of 153 requests wrote whitespace or carriage returns until max_tokens, with the keys in alphabetical order 16 of 16 replayed requests were valid and in the schema's order, with at most 3 whitespace characters in a row (September 27)
Use 8-bit or 4-bit KV cache on the 27B with a long answer (#526) Every logit came back as NaN, and the request failed with HTTP 500 An 18,893-token answer that counts to 4,000 finished correctly at both settings (September 27)
Continue a long 27B agent session on a 48 GB Mac (#499) A 99,355-token turn produced nothing in 1,224 s before OpenCode gave up Refused before its prompt is read, with a 507 that says why (tests with the reporter's numbers)
Send a 24,026-token prompt to Gemma 4 31B 92.08 GiB peak, 173.5 s to the first token 31.1 GiB peak, 45.1 s, and a 33,026-token prompt fits in 37.75 GiB (September 27)
Use OpenCode with a context window under 64K tokens Every answer was capped at OpenCode's reply reserve: 16,384 tokens on a 32,768-token window Answers are no longer capped by the reserve (tests)
Pick a model folder in mtplx start (#573) The next start offered a 20.7 GB download of the default instead The folder loads on every start (tests)

Rows dated October 2 come from a build of this release running in the app on an M5 Max with 128 GB, with Flash-Next Optimized Speed at the default 90 GiB limit and Pi. The compaction ran on a copy of that session while the Mac was busy with other work, so its times are not quoted here. The #581 row comes from the reporter's script against mtplx serve with Qwen 3.8 27B Optimized Speed on the same Mac, five runs on each version. Rows dated September 27 come from builds of the same fixes on the same Mac, and rows marked tests come from the tests of each fix. The 2.12.0 column comes from bug reports, from 2.12.0's code and from runs on the same Mac; the Gemma 4 row's first column was measured on a September 27 build whose Gemma prompt processing was still 2.12.0's.

Memory: the guard, rewritten

These changes apply to every model and every Mac. They were measured on an M5 Max with 128 GB. The limits below scale with each Mac's memory, but no Mac with 8 to 64 GB has run this release.

  • Every request is priced before its prompt is read. 2.12.0 checked only requests with 4,096 or more uncached tokens, and it charged those tokens once plus a flat 3 GiB, so a short turn on a long conversation was never checked. Now every request is priced at its largest moment: the end of prompt processing (including the copy a restore makes), the move to paged memory, and the start of the answer. On a 48 GB Mac with the 27B, the turn from #499 is priced at 39.47 GiB and refused before its prompt is read, where 2.12.0 admitted it and reached 38.5 GiB against a 36 GiB limit. On a 64 GB Mac with 4-bit KV cache (#525), a 250K-token prompt was projected at 7.6 GiB and needed 20.1 GiB; the price now covers the full-width rows that exist while the 4-bit pages are filled.

  • Idle conversations go before anything is refused. When memory is short, the engine first gives back what it holds for itself: its buffer pool, then the RAM copies of conversations that are not generating, oldest first and those already saved to SSD first. A conversation that is generating is never touched, and a request that arrives for a conversation being released waits for it instead of getting a 409. A new request, a compaction included, no longer waits behind an idle conversation's memory. A 507 that remains names what holds the memory and when a retry can work.

  • The engine reads what macOS can actually hand out. 2.12.0 read macOS's memory level, which counts memory that macOS can free only by compressing other apps as available. The engine now reads free, purgeable and file-backed pages, checks them before every chunk of a prompt, and stops a prompt with a 507 instead of pushing the Mac into compression. The stop floor is the largest of 1 GiB, 2.5 percent of RAM and a sixteenth of the memory macOS has wired, and the shed floor is twice that: 6.2 to 6.4 GB and 12.4 to 12.8 GB in the October 2 session. A prompt also stops when macOS has compressed another eighth of RAM during a run of requests (16 GiB on a 128 GB Mac), or holds a quarter of RAM compressed while it is still losing ground. The early compression that macOS 27 does on a healthy Mac does not count. On September 29, 2.12.0's 130,048-token cold Flash-Next prompt drove 27 GB into compressed memory until the test's watchdog stopped it, and a build of this release (0e0bb014) served the same prompt with no swap and 15 MB of compressor growth. That prompt has not been run again on the final build.

  • Memory outside the GPU allocator now counts. 2.12.0 forgave up to the larger of 8 GiB and RAM minus the system reserve minus the limit, so a lower limit raised the allowance by the same amount and the ceiling never moved: 112 GiB on a 128 GB Mac at a 96, 90 or 88 GiB limit. The allowance is now a sixteenth of RAM, between 1 and 8 GiB (8 GiB from 128 GB up), memory past it is charged, and an explicit limit caps it at a twelfth of the limit. A lower limit now lowers the whole ceiling, and on October 2 the process ceiling was 98 GiB at the 90 GiB default.

  • The buffer pool goes back to macOS. MLX keeps freed GPU buffers for reuse, up to 8 GiB on a 128 GB Mac, and 2.12.0 kept them through the client's whole think-and-tool time: 0.53 GiB at the median and 3.5 GiB at p90 over 4,104 logged requests. The pool now goes back after every completed request and after a second with no work. While work runs it keeps its full size, because a 2 GiB bound cost 6 to 10 percent of the prompt speed at 16K and 64K tokens. MTPLX_CLEAR_CACHE_AFTER_REQUEST=off keeps the pool.

  • The prompt chunk narrows before anyone's memory is taken. When the Mac is short, the engine first clears its own buffer pool, then runs Flash-Next's prompt in 2,048-token chunks instead of 4,096, and only then releases conversations. The chunk width is chosen again after memory is freed, so a prompt that fits once idle state is released keeps its 4,096-token chunks.

  • The price covers what each model really allocates. That includes a quantized KV snapshot at full width (a 250K-token 27B snapshot is 15.3 GiB, not the 4.9 GiB of its 4-bit pages), the routed experts of MoE layers, the draft head's history (4,096 bytes a token on the 27B) and Gemma 4's own prompt pass.

  • Pressure trims reach the session cache (#525). On a 64 GB Mac with the 27B, 421 trims in a row under macOS memory pressure evicted nothing, because their target was half of the cache's budget and the cache held less than that. A trim now halves what the cache holds, and it spares conversations that are generating.

  • A guard step that fails is reported. If one of the guard's own steps raises, the request still runs, so the guard never costs you an answer, but /health and the dashboard stream now show memory_guard.guard_degraded with the error.

The default memory limit on 128 GB Macs is 90 GiB

From 128 GB up, the engine now leaves at least 38 GiB outside its allocator for macOS and your other apps, so a 128 GB Mac plans with 90 GiB instead of 96. Macs under 128 GB, and from 192 GB up, keep the 75 percent rule. On September 27, with 16 GB of other apps open, a 74,000-token compaction-style request was refused with a 507 at 96 GiB to keep the Mac safe, and served at 90 GiB (build aa8ec218). A model whose resident floor is between 90 and 96 GiB keeps 96 GiB, and one above 96 GiB keeps everything outside the system reserve, as before.

What it costs:

  • On 128 GB Macs with M1 to M4 chips that run Flash-Next without its native prompt kernel, which pip and Homebrew installs do not carry, Flash-Next Optimized Speed plans 65,536 tokens instead of 114,688. From the app, which carries the kernel, the same Macs plan 262,144 tokens when the kernel passes its startup check; that check has not run on M1 to M4 hardware yet. M5 Macs plan 262,144 tokens either way.
  • Qwen 3.8 27B Optimized Quality's session cache under a full 262,144-token window is 42.5 GiB instead of 48 GiB.

--memory-limit 96G on mtplx serve, MTPLX_MEMORY_LIMIT_BYTES=96G, or 96 in the app's Settings, Memory, Memory limit field restores the old limit. --memory-limit max (#548) uses everything outside macOS's own reserve, 112 GiB on a 128 GB Mac, for a Mac that runs nothing but the server; the whole-Mac checks still apply. On Macs with 32 GB or less, max is the same as the default.

Long sessions keep their cache

  • Screenshots no longer cost the conversation. The session cache now matches an image conversation by the content of its images, never restores to a point inside an image, and keeps restore points between and after images. A Pi turn that changes text after a screenshot resumes at the change instead of reading the conversation again. A request with different pixels in the same place still never reuses the old image's state.

  • Flash-Next no longer copies a conversation to verify it. The compiled verifier now adopts the buffers the prompt was read into instead of copying them into a padded buffer, and the session cache keeps a reference to that live cache instead of a second copy. When the buffers must grow, they grow in place one layer at a time, and a turn that cannot afford the growth runs on the eager verifier for that turn only, instead of keeping every later turn eager. This removes those copies by design; how much memory it saves has not been measured on its own. MTPLX_ONE_COPY=0 restores the copying store.

  • An answer that cannot grow ends cleanly. If memory does not allow the cache to grow during an answer, the answer now stops between two rounds with every streamed token kept, finish_reason length, memory_stop in the stats and a plain message, and the next turn reuses the whole conversation. Before its prompt is read, a request's longest answer is priced too: when only a shorter answer fits, the request runs with that limit and says so if it reaches it. In the October 2 session that limit applied from 112K tokens on (28,014 tokens at 205K), and no answer reached it. The stop itself has so far run only in tests on small models.

  • Stop, cancel and resend keep the conversation. A cancelled answer now hands its cache back in one step, so a failure halfway can no longer leave a half-converted conversation that looks reusable. A prompt that stops part-way, for example when the client cancels a long re-read, gives the conversation back at the point it was restored from. In 2.12.0, a repeated prompt that restored a whole saved conversation resumed one position short, so its first answer token overwrote the last prompt token; it now resumes at the right position.

  • A conversation's own saved state survives the memory check. The entry a conversation will restore from is protected in every release step, and its pending SSD save is never cancelled. When memory is short, that entry moves to SSD first. When nothing can put it on disk, it is released as the last step before a refusal, and the conversation reads it again instead of being refused.

  • The turn after a compaction gets its memory. A refused request now frees its memory at once, so a retry does not meet it (4.95 GB stayed active in one case). The SSD cache reads a saved conversation's description before its tensors and skips one that cannot be restored, instead of reading gigabytes for nothing (4.17 GB in one case), and it tries the next candidate. Before refusing, the check during prompt processing releases this conversation's own unusable entries.

  • More restore points survive. Recurrent restore points are now kept within a byte budget on a fixed 8,192-token grid, instead of being thinned after every capture. A 16,371-token match that restored at 4,096 now restores at 8,192.

  • Saving a turn no longer pushes out the entry the next turn needs (#554 by @jvmenen). Reading a conversation's own saved prompt to re-render its history no longer counts as use, so it cannot outrank a sibling entry, such as a retry that diverged early.

SSD session cache

  • A new save no longer deletes the old one first. Both writers now write the new data beside the existing entries, install the new entry and only then retire old entries. A save that does not fit is skipped, the saved copy stays, and the reason is recorded. Before, a save could delete the only saved copy first: a 131,735-token save evicted it two minutes before its own entry landed, and a request in between found nothing.

  • The SSD cache holds two copies of your largest conversation. Its size is normally a quarter of the space it can use. It now grows to two copies of the largest conversation when the disk has room above a 10 GiB floor, and never exceeds the configured cap. Before, with 25 GiB free the cap was 6.25 GiB, room for one copy and not its replacement.

  • Queued SSD saves can no longer hold memory the cache already gave back (#536 by @vahako). With four agent clients that each open new sessions, memory climbed from 83.4 to 104.1 GB in four minutes on a 128 GB M5 Max until requests failed with 507 (the author's measurement on 2.12.0). Saves whose entry the cache has already dropped now count against a budget of a thirty-second of RAM (4 GiB on 128 GB, at least 512 MiB), and the oldest are cancelled above it. A save whose entry is still in RAM costs nothing extra and is never cancelled, so a long main session keeps its disk copy while subagent sessions come and go. MTPLX_PERSISTENCE_MAX_PENDING_BYTES sets the budget.

Coding agents

  • Pi and OpenCode are set up with the window the engine serves. /health now publishes execution_window, the context window the engine can actually serve on this Mac. The app syncs Pi's and OpenCode's settings to it once the engine answers, and mtplx start pi, mtplx start opencode and mtplx connect opencode do the same, also against a server that is already running. Before, the CLI wrote the window from the model's metadata or --context-window before the server started (262,144 tokens against a served 32,768 in one case), and Pi planned answers and compaction for a window the server could not serve. Pi's answer ceiling stays the whole window, as in 2.12.0, and a window pair you set yourself is never changed.

  • OpenCode answers are no longer capped on small windows. OpenCode sends its reply reserve, half the window up to 32,000 tokens, as the answer limit of every request, and 2.12.0 removed it only when it was exactly 32,000. On windows under 64K tokens the reserve reached the server as every answer's cap: 16,384 tokens on a 32,768-token window and 10,240 on 20,480. Both generated plugins now remove the reserve at every window size, and a limit you chose still passes through.

  • An answer cap you type for OpenCode reaches the server. mtplx start opencode --max-response-tokens N against a server that was already running lost N: it was clamped into OpenCode's reply reserve and then removed as OpenCode's own default. A typed cap now reaches the server whole, also above OpenCode's 32,000 ceiling.

  • Attaching to a server that serves another model works. When mtplx start pi or mtplx start opencode finds a server already running another model, it now sets the client up for the model that server serves, with its window, image support and reasoning settings. Before, Pi opened on a model its config did not list, and OpenCode stayed on the model the CLI had picked.

  • Client configs are read the way the clients read them. The app read Pi's models.json and opencode.json as strict JSON and moved a file it could not parse aside, so a comment or a trailing comma lost every other provider and your own window settings on the next sync. Both files are now read with comments, trailing commas and a byte order mark, as the clients do. A file that still does not parse is left exactly as it is, and the app says which file and why.

  • A stream ends at its last token (adapted from #557 by @jvmenen). For a session the client names, the last chunk and [DONE] no longer wait for the session's save, which took 18 ms at the median and 160 ms at p90 in our logs and up to 30 s behind another client's prompt. Only the next request of the same session waits for that save, so it reads exactly the state it read before, and other sessions never wait. A turn sent the instant the previous stream ends now waits for that save instead of the stream. If the model makes no progress for the stall deadline (300 s), that next request gets a 503 (session_commit_stalled) instead of reading the session early. MTPLX_STREAM_TERMINAL_FRAME_BEFORE_COMMIT=0 restores the old order.

  • Finished answers are no longer re-generated. The server re-ran a whole turn when an answer after tool results ended with text its pattern read as a promise to act, such as Say "A" or "B" and I'll run it., and threw away finished answers (82 s of them in one session). That retry is removed. The two re-generations that remain build their prompt from the tokens the request was actually served, so they continue the saved conversation instead of reading it again; when no exact prompt exists, the retry stands down and the first answer is delivered.

  • A long tool call is no longer cut off. The stream cancelled any tool call that had buffered 2,048 tokens and 30 s outside a known parameter, and told the client the tool call was malformed. That cut long valid writes, so the guard is now off by default (see Experimental).

  • Fewer and better retries (#539, #540, #552 and #553 by @jvmenen). A well-formed tool call whose argument is one short value is no longer retried as broken tool markup; in @jvmenen's agent loop on Qwen 3.6 35B-A3B, long tool turns reached their first token in 10 to 20 s instead of 25 to 87 s. A blank answer at temperature 0 is no longer regenerated three more times, because a greedy retry repeats itself: a one-token completion took 265 ms instead of 1,061 ms, and usage.completion_tokens now describes the returned attempt. The reasoning repair now continues a retried turn from the retry's own prompt, and the stream's stats count every pass. MTPLX_AGENT_REWRITES=off now also turns off the two retries that inject a steering message.

  • Parallel tool calls use the model's own tool format (#270 by @shiftedx). Pi and Hermes requests that set parallel_tool_calls: true now use the tokenizer's native tool template, which writes sibling calls in one turn more reliably. Requests that leave it out or set it to false keep the existing format, so the session's prompt stays stable. Native tool prompts now keep the client's own system instructions, and a forced tool_choice is honored in that format.

  • Agent turns tokenize only what is new (#533, #541 and #542 by @jvmenen). Each tool-call turn used to tokenize the whole conversation again on the CPU before the prompt was processed. Turns are now encoded and remembered one by one, with token ids identical to a single encode, and plain chat is cut into turns only when the tokenizer is proven to give the same ids. Measured by @jvmenen on an M5 Pro with Qwen 3.6 35B-A3B: an 82K-token agent transcript took 57 ms instead of 107 ms to encode, a short follow-up to an 80K-token chat 1.3 ms instead of 66 ms, and a check for Gemma 4 on every chat request takes about 15 microseconds instead of building the whole vocabulary (about 47 ms).

Fixes

  • Flash-Next prompts on M5 Macs no longer read leftover memory in the expert step. MLX 0.32.2's tensor-unit kernel for expert-sorted quantized matrix multiplies skipped rows when one call routed more than 32,767 rows and the count was not a multiple of its tile, and those rows kept whatever the memory held. Flash-Next sends each token to ten experts, so a prompt chunk of 3,277 to 4,095 tokens could cross that limit: in 2.12.0, 794 of every 4,096 cold prompt lengths above 3,277 tokens (19.4 percent), and warm suffixes over 512 tokens by the same rule. Such a call is now padded and run so that every real row is computed exactly as in a correct call. The guard covers every MoE model MTPLX loads; M1 to M4 Macs and MLX 0.32.3 were never affected.

  • Structured output no longer loops (#547). A strict json_schema request could write most of the object and then emit whitespace until max_tokens: the reported case ended in two spaces and 640 carriage returns, with the keys in alphabetical order. The grammar now keeps the client's key order and allows at most one space, or one or two newlines plus up to 20 spaces or tabs, between JSON tokens, and never a carriage return. A client that sets its own whitespace rules in x-guidance gets exactly those. llguidance 1.8 or later is now required, so every install gets the same layout.

  • Screenshots after tool calls work in streaming clients (#581). A streaming request whose last user message carried an image after a tool call could fail with "vision requests must not use the session bank without content-keyed ids". That is the request OpenCode sends after a screenshot tool, and non-streaming requests were not affected. With the reporter's script on Qwen 3.8 27B Optimized Speed, 3 of 5 streaming runs failed on 2.12.0 and 5 of 5 passed on 2.12.1.

  • Constrained drafting follows the right distribution. With a grammar and top-k or top-p sampling, drafted tokens were checked against rows the grammar had not masked, which over-weighted some legal tokens compared with plain decoding. Every drafted row is now masked at its own grammar position, so constrained output follows the same distribution as plain decoding, and greedy constrained output commits the same tokens as before with more drafts accepted. The session cache keeps the model's own unmasked logits, so a later request on the same prompt is not bound by the earlier request's grammar.

  • 4-bit and 8-bit KV cache no longer turn into NaN (#526). With KV quantization on, the compiled verifier's paged cache was sized once, for the prompt plus 16,384 tokens, and never grew, so a longer answer wrote across heads and past its buffers. The cache now grows before any write, the growth is checked against the Mac's memory and refused with a 507 before a row moves, In our reproduction on the 27B, an answer counting to 4,000 failed with NaN on 2.12.0 at 8-bit KV and now finishes correctly at 8-bit and 4-bit; no M3 Mac, where it was reported, has run the fix. The cost: below the old capacity, the 27B decodes about 3 percent slower at 8-bit KV (57.42 against 59.25 tok/s on the fix's own build, alternating boots, September 27).

  • FP16 builds keep attention's partial sums in float32. On FP16 builds, used on M1 and M2 Macs, a block of large attention scores could store infinity before its division and turn every output of that call into NaN. These partial sums are now kept in float32 for FP16 queries in every split attention kernel. BF16 outputs are bit-identical to 2.12.0 (108 of 108 outputs across 11 kernel families), and FP16 builds use twice the scratch for these sums (25 to 50 MB per layer call for the 27B at 32K tokens). No M1 or M2 Mac has run this fix.

  • The GPU event leak behind "Failed to create Metal shared event" is fixed (#544). Flash-Next's sparse attention index and the draft head's history were written lazily and never read in some rounds, and each unread write kept one Metal event alive. A non-streaming request saves its final state in the session cache, which kept those events until the entry left, so a long non-streaming run eventually ran out. The writes are now evaluated with the step that makes them. In an A/B on September 27 the live event count stayed at 6 to 13 with the fix against about 2,900 on 2.12.0, and on a build of this release on October 2 it read 1 to 9 through three 4,096-token non-streaming answers. A run past the reporter's 140K generated tokens has not been repeated.

  • Host memory grows more slowly in long sessions (#546, partly fixed). MLX keeps every compiled version of the shared verify program, and a growing session added one almost every turn. A shared program is now replaced after 16 versions (MTPLX_COMPILED_VERIFY_TRACES_PER_PROGRAM), and long Flash-Next conversations size the verifier in 8,192-token steps, so one program serves many turns. Part of each replaced version stays inside MLX, so the growth is slower and capped per program rather than gone, and a restart still returns it. In the October 2 session, host memory outside the GPU rose from 2.72 to 3.86 GB between 158K and 205K tokens and fell to about 2.7 GB once that conversation's cache was released.

  • Draft heads with fused MoE experts load their experts (#574 by @jvmenen). Qwen 3.5 and 3.6 MoE packs built from Qwen's official BF16 release store the draft head's experts as fused tensors, and the loader dropped them, so the experts kept their random start values. They now load. On @jvmenen's Forge pack of Qwen 3.6 35B-A3B, the draft acceptance at depth 2 went from 0.69 and 0.37 to 0.95 and 0.85, the same as the head saved the other way.

  • Batched requests keep their padding out of shorter prompts (from #522 by @freddyhaddad). With --scheduler-mode ar_batch, a Qwen 3.5 or 3.6 hybrid request processed beside a longer prompt ran the padding tokens through its recurrent state and decoded from a drifted state. The default serial scheduler was never affected.

  • Streamed replies no longer start with a blank line (#531 by @pierre427). When the blank line after a model's closing tag arrived in a later chunk, it streamed as the first visible text, which Gemma 4 replies showed in streamed clients only.

  • Gemma 4 sees its own tool calls in the format it is asked to use (#571 by @jvmenen). Gemma 4 packs ship no chat template, so the tool history was flattened into text while the prompt asked for structured tool calls. The model copied the flattened form back, left calls unfinished more often, and each tool turn stopped matching the saved prompt. The history now stays structured, and earlier turns render with the same bytes the model generated.

  • Flash-Next's two verifiers agree at near ties. The compiled verifier divided its block scores by a constant that MLX writes into compiled code with 7 significant digits, so at a near tie it could pick a different attention block than the eager verifier. Four gates in Flash-Next also rounded one value differently in the two verifiers. The divisor is now read from memory, and the eager verifier rounds those gates the way the compiled one does, on every Mac.

  • A non-finite answer says where it went wrong. When a request fails with non-finite logits, the log now names the cache, the attention route, the offset and the capacity of the first layer involved, per request. MTPLX_KV_ATTENTION_TRACE=nonfinite checks every call and prints only the bad ones. The /health text and the --kv-quant help now describe the KV quantization routes that really run.

Gemma 4

  • Long prompts fit. Gemma 4 processed every new prompt token in one pass, and its attention scores grew with the square of the prompt. On the 31B pack, cold prompts of 6,026, 12,026 and 24,026 tokens peaked at 25.16, 43.02 and 92.08 GiB, with 12.9, 41.4 and 173.5 s to the first token. The prompt now runs in chunks: the same prompts peaked at 21.54, 25.48 and 31.1 GiB with 11.7, 20.3 and 45.1 s to the first token, and a 33,026-token prompt was served in 37.75 GiB (September 27). The memory check prices Gemma's own prompt pass from its measured peaks, and the request can be stopped between chunks.

  • Its output changes by rounding. Chunked attention adds the same terms in a different order, so Gemma 4's text now differs from 2.12.0's: at 4K tokens, from the eleventh word on. Both read correctly. MTPLX_GEMMA4_PREFILL_CHUNK_TOKENS=whole restores the single pass, whose first turn matches 2.12.0.

  • A trade-off. A saved Gemma 4 prompt can now be restored after an edit only when the edit falls within its last 2,047 tokens; an earlier edit processes the prompt again.

  • Prompt scoring works on Gemma 4 (#551 by @jvmenen). /v1/completions with echo, logprobs and max_tokens: 0 answered HTTP 500 on Gemma 4. It now scores within the same prompt width and memory policy as generation, and is refused with a 507 under memory pressure.

  • First-token log-probabilities (below), the native tool prompt fixes and the blank-line fix also apply to Gemma 4.

Speed

This release was not timed end to end against 2.12.0, so it makes no overall speed claim. The figures below are for each kernel or module alone, measured on an M5 Max with synthetic weights and MLX 0.32.2. Every change gives the same bits as the code it replaces, and each new kernel compares itself with the stock operations on first use and falls back to them, with a printed reason, on any difference.

Change Who gets it Measured alone
Flash-Next prompt processing reads each routed token row in place and fuses SiLU into the gate and up step M5 Macs 10.0 ms to 7.65 ms per layer at a 4,096-token chunk
The same kernel for Qwen 3.5 and 3.6 MoE models such as the 35B-A3B, whole expert module M5 Macs 1.19 to 1.24 times as fast at 512 to 4,096-token chunks
Flash-Next block selection at verify widths, 12 layers in a chain Flash-Next's compiled verifier 0.843 ms to 0.590 ms per verify step at 4K, 1.563 ms to 1.339 ms at 64K
Flash-Next pooled index key row, 12 layers in a chain Flash-Next's compiled verifier 1.380 ms to 0.759 ms per verify step at 4K, 1.061 ms to 0.653 ms at 64K
Flash-Next expert step at verify width, with the residual tail in the same dispatch Flash-Next Optimized Speed 275.4 µs to 257.0 µs per layer
  • Flash-Next hyper-connection reads at verify widths run as three kernels instead of about eleven, on M3, M4 and M5 Macs with one die. M1, M2 and Ultra chips keep the stock path. MTPLX_QWEN4_HC_VERIFY_READ=0 turns it off.
  • The draft head does less work during prompt processing on Qwen 3.5, 3.6 and 3.8 packs (#555 by @jvmenen). Its history append now evaluates only the cache it writes, with the same cache bits, on every Mac. Flash-Next already did this in 2.12.0. MTPLX_MTP_HISTORY_CACHE_ONLY=0 turns it off.
  • Sampled copy rounds read the GPU once per 8 rows instead of once per token: a fully accepted 24-token block takes 3 reads instead of 24, with the same tokens for every seed.
  • Long Flash-Next conversations re-compile the verifier less often. Its buffers are sized in 8,192-token steps, which cut verify re-traces by 95 percent at 16K to 64K tokens. MTPLX_QWEN4_FIXED_M4_CAPACITY_BUCKET sets the step, and 0 turns it off.
  • Flash-Next keeps its wide prompt chunks after memory is freed. The 4,096-token chunk is chosen after the memory check releases idle state, not before.
  • Prompt scoring picks its top results without a full-vocabulary pass (#532 by @jvmenen): 3.4 to 5.2 ms per chunk on the GPU instead of 47 ms, on synthetic data on an M5 Pro, with bit-identical values.

Exactness

The release is lossless by default. The new kernels and the new memory and cache paths give the same tokens and logits as the code they replace, and each new kernel compares itself with the stock operations on its first call and falls back to them on any difference. Output differs from 2.12.0 only where a fix requires it:

  • Gemma 4 processes its prompt in chunks, which changes its text by rounding (see Gemma 4). MTPLX_GEMMA4_PREFILL_CHUNK_TOKENS=whole restores the single pass.
  • Flash-Next's compiled verifier now reads its block-score divisor from memory instead of a 7-digit copy, so at a near tie it picks the same attention block as the eager verifier; such a tie can pick a different block than 2.12.0 did. Flash-Next's eager verifier now rounds four gates the way the compiled one does.
  • Structured output keeps the client's key order and bounded whitespace (#547), and with top-k or top-p sampling, constrained drafting now follows the same distribution as plain decoding.
  • Draft heads with fused MoE experts now load their experts (#574). The drafts change, and the sampled output distribution does not.
  • Flash-Next prompts on M5 Macs that crossed MLX 0.32.2's row limit now compute every expert row instead of reading leftover memory.
  • 4-bit and 8-bit KV cache answers that outgrew their cache now continue correctly instead of failing with NaN (#526).
  • A repeated prompt that restored a whole saved conversation now resumes at the right position.
  • pip and Homebrew installs now get MLX 0.32.2, the version the app bundles and every check ran on. An install that had picked up MLX 0.32.3 produced different Flash-Next output.

Two changes end an answer early without changing its tokens: an answer whose cache cannot grow stops between rounds, and an answer is limited to what memory can hold. Both say so in the response.

The app

  • The badge no longer sticks on Degraded (#528). A second start against the engine the app already runs, for example a reopened main window, the Start command, the benchmark or Hermes, showed Degraded with a message that MTPLX was already running, until you quit the app. The app now reconnects to its own engine and shows Running, and Refresh and the live stats run the same check. The badge separates the engine from the live stats: a running engine whose stats stream dropped reads "Running · reconnecting live stats".
  • Starting, stopping and restarting are safer. Closing the main window during a model load no longer kills the load. Start or Restart after Stop launches instead of joining the start that Stop cancelled, and Stop while a start is still preparing means nothing launches. Requests go to the port the engine really listens on, also after a fallback port. A configuration change that fails before its restart leaves the engine running and shows its own notice. Stop signals only processes that carry this app's launch id, a server you started in a terminal on the same port is named and never stopped, the benchmark runs only on the app's own engine, and requests keep using the API key the running engine started with.
  • The app says why a prompt is read again. While a prompt is read again, the Live gauge and the Activity and Requests rows say why (the history changed at a token, a changed screenshot, a new conversation, state freed for memory, and so on), where the state resumes, what the SSD restored and how long the server estimates it will take. Every cache miss the server reports now has a plain sentence, and estimates are shown only after a prompt speed was measured.
  • A banner warns when the SSD cache is short of disk: low when two copies of the largest conversation do not fit above the 10 GiB floor, and full when saves have stopped.
  • The live gauge shows the current phase's speed. It used to divide every token by the time since the first one, so a slow reasoning phase dragged the number down for the whole answer. It now shows the rate of the current phase (reasoning, answer or tool call), and the finished request keeps its overall rate.
  • Dashboard axes fit their labels (#572 by @jvmenen). The decode and prefill charts clipped labels such as "52 tok/s".
  • New text in the app is translated into all 13 languages.

Server, CLI and API

  • Log-probabilities for the first generated token (#530 and #543 by @jvmenen). /v1/completions accepts logprobs: K with max_tokens: 1, and /v1/chat/completions accepts logprobs: true with top_logprobs: K. The values are the model's raw distribution for the first token, before temperature, penalties, grammar and steering, and a full-prompt cache hit uses the restored logits. A classifier can read the next-token distribution through the normal generation path with prefix reuse: in @jvmenen's test on an M5 Pro, it gave the same top label as prompt scoring on 228 of 240 prompts, with a median of 486 ms against 562 ms. Streaming, stop sequences, max_tokens other than 1 and K above MTPLX_PROMPT_LOGPROBS_MAX (128 by default) return 400, and /v1/messages does not offer it. The chat entries report each token's own bytes, or null when the tokenizer's decoder cannot give them.
  • Every response says where its time to the first token went. mtplx_stats.ttft_spans.exclusive_s splits the time from the request's arrival to the first visible delta into named spans, such as parsing, encoding, waiting for the previous turn's save, the queue and what it waited behind, the memory check and the engine. attempts, discarded_attempt_wall_s and retry_path name any generation a repair threw away, and /health shows the median and p90 of each span. ttft_s now starts at the request's arrival, so it includes reading and parsing the body.
  • Thermal pressure is recorded. /health, every live sample and every request receipt now carry macOS's thermal pressure (thermal_pressure and the worst level during the request, thermal_pressure_max), read from a cached sampler so decoding makes no extra system call.
  • /health reports the served window and the guard's health: execution_window (above) and memory_guard.guard_degraded. session_bank.last_ram_miss_reason says why RAM could not serve a shared prompt start (#534 by @jvmenen), and each request's metrics row records what the Mac's memory did during its prompt.
  • mtplx start keeps the model you picked (#573). A model folder picked in the "Local folder" step was swapped for this Mac's default repository on the next start, and Yes started a download of a model already on disk. Setup now records where a pick came from. A folder loads on every later start, a catalog pick that is not this Mac's default stays as picked, and only the verified default follows when the default moves (now also Flash-Next on 256 GB Macs). A config.toml model folder loads as configured, whatever it is named, and a missing folder is named instead of replaced by a download.
  • Dependencies. pip and Homebrew install exactly MLX 0.32.2, the version the app bundles; the requirement allowed 0.32.3, whose Flash-Next output differs and which none of this release's checks ran on. llguidance 1.8 or later is required (#547). transformers may be up to 5.16, the releases checked against the packs' tokenizers (#538 by @dependabot).

Experimental, off by default

  • MTPLX_SSD_INCREMENTAL_ENCODE=1 (#556 by @jvmenen). The SSD cache saves only the parts of a conversation that changed since its last save, instead of encoding and hashing all of it on every turn. It is off because the fingerprint that decides a block is unchanged is not a content identity: flipping the signs of any two float32 values leaves it the same, so a changed block could reuse its predecessor's data and a later restore would bring back the wrong state. The full save stays the default until the key cannot collide.
  • MTPLX_FIXED_M4_COPY_WINDOWS=1. Flash-Next's context-copy rounds replay on the compiled verifier instead of the eager path, bit-identical to the eager window in tests. It is off because it has not been checked on the full model yet, so no speed is claimed for it.
  • MTPLX_THINKING_GUARD_SCOPE=all (#545 by @bpmforge). The reasoning budget guard, which you turn on with --agent-thinking-budget or MTPLX_THINKING_BUDGET, covered only requests that declare tools. With this setting it also covers plain chat with thinking on. It is off because it can end a model's reasoning early, and generation stops ship off.
  • MTPLX_STREAM_HIDDEN_TOOL_GUARD=on. The guard that cancels a tool call after 2,048 hidden tokens and 30 s is now opt-in, because it cancelled long valid tool calls. MTPLX_STREAM_HIDDEN_TOOL_GUARD_TOKENS and MTPLX_STREAM_HIDDEN_TOOL_GUARD_S tune it once it is on.
  • Diagnostics. MTPLX_KV_ATTENTION_TRACE=1 prints a line for every full-attention call, and nonfinite prints only calls that produced non-finite values. MTPLX_PREFILL_GUARD_TRACE=1 prints every memory check during a prompt.

For developers

  • Tests no longer read your own settings (#535 by @jvmenen). The suite read ~/.mtplx/config.toml, so a real config changed dry runs and sent Forge builds into your model cache, and one route test failed on Macs with less than 85.3 GiB. Tests now point every client config they can write (Pi, OpenCode, OpenCode Desktop and the mtplx start state) at scratch files.
  • Exactness tests are stricter. The compiled-verifier parity tests also run with MLX_ENABLE_TF32=0 and under the M1 to M4 rehearsal switch, their toy weights are evaluated the way a loaded checkpoint's are, and bit-equality checks compare float32 and float16 by their bit patterns, so a sign-of-zero difference fails.
  • New kernels build from the installed MLX's own headers and keep the stock path, with one printed reason, when a header cannot be read.

Validated

Check Result
Python suite on the final code 11,264 passed, 37 skipped, 1 expected failure, 0 failures
M1 to M4 rehearsal on the M5 763 passed, 200 skipped
Swift app suite 1,104 tests, 2 skipped, 0 failures
The app on the final code with Qwen 3.8 27B Optimized Speed, October 2 24 checks through the app's own engine passed: streaming and non-streaming answers, cancel and resume, cached follow-ups, images, strict JSON, tool calls, the Chat Completions, Responses and Anthropic Messages formats, two queued requests and malformed input. Swap stayed at zero
A Pi coding session in the app on Flash-Next Optimized Speed, October 2 87 requests, 12 screenshots, up to 204,981 tokens; no errors, no 507s and no memory stops; 86 of 87 turns from RAM (the other turn hit a defect fixed before release) and 99.2% of prompt tokens reused; median first token 1.07 s; 50,769 of 50,769 verify steps compiled; peak GPU memory 94.2 GB under the 96.6 GB limit
A real Pi compaction of that session, October 2 Two summary requests and the turns after them served, with no 507 and swap flat at 2.08 GB; the lowest free memory was 259 MB
The #544 event count on a build of this release, October 2 1 to 9 live events through three 4,096-token non-streaming answers
The memory guard at 96 and 90 GiB with 16 GB of other apps, September 27 At 96 GiB a 74,000-token compaction-style request was refused safely, and at 90 GiB it was served
Answers whose cache cannot grow, and cancels during growth 42 targeted and 1,005 packaged-wheel tests on the recovery change, then 143 and 173 targeted tests on the two commits after it

Known issues

  • 32 GB Macs get an 8,192-token window on the default pick. The first suggestion there is Qwen 3.8 27B Optimized Speed, which leaves room for 8,192 tokens: enough for chat, too small for a coding agent. On the same Mac, Ternary Bonsai 2 27B plans 192,512 tokens and Qwen 3.8 27B Bare Speed about 77,824.
  • 48 GB Macs on the 27B (#499). Pi is told 204,800 tokens, but a warm turn near 100K tokens is refused with a 507, because the 27B still keeps two copies of a restored conversation; the single live copy is Flash-Next only. It is a clean refusal now instead of a swap stall. On the 27B, contexts past 32,768 tokens verify on the slower eager path, as in 2.12.0.
  • Flash-Next on a 128 GB M3 or M4 Max plans 262,144 tokens from the app but 65,536 from pip or Homebrew, which do not carry the native prompt kernel.
  • Older Macs. On M2 and M3 Macs the 27B can still hit macOS's GPU watchdog past about 100K tokens (#422), and Flash-Next on an M2 Max can stall at about 1.5 s per verify round (#400). 8 GB M1 and M2 Macs still get no recommended model (#521).
  • Measured on one Mac. Every real-model check ran on an M5 Max with 128 GB. M1 to M4 were rehearsed on the M5, which checks the code paths and results but not an older GPU's compiler, speed or watchdog. No Mac with 8 to 64 GB, and no Ultra, has run this release. On 8 GB Macs the 4B is priced with the 27B's fixed scratch figure, so some requests that would fit may be refused.
  • Compaction reads the whole conversation again. Pi's compaction sends a new, rewritten prompt that cannot reuse the saved conversation. In the October 2 test, on a Mac busy with other work, compacting a 205K-token session took 410 s. After a compaction, the app's explanation can blame memory pressure for a prompt that simply does not share the saved history.
  • A compaction can still be refused when the rest of the Mac is short. With a heavy desktop, a compaction can be stopped about 20 s into its prompt with a 507, which Pi shows as an error. The Mac stays safe.
  • The re-read estimate is optimistic on long prompts. For a 204,817-token re-read the app estimated 3 min 55 s, and it took 7 min 43 s on a busy Mac.
  • Memory checks cover prompt processing. Reading a saved conversation back from SSD before a prompt, and the eager verifier's own cache growth during an answer, are not checked yet. --allow-swap (or the app's swap setting) turns off the checks during prompt processing and the refusals before it.
  • Host memory still grows in very long sessions (#546), more slowly than in 2.12.0. A restart returns it.
  • Long non-streaming runs (#544): a run past 140K generated tokens has not been repeated on this release.
  • Heat limits long sessions. In the October 2 session, macOS reported heavy thermal pressure in 70 of 87 requests with the fans at maximum. Requests now record the thermal state.
  • mtplx serve ignores MTPLX_SSD_SESSION_CACHE=off, because it passes its own --ssd-session-cache on. Pass --ssd-session-cache off instead.
  • Gemma 4 prompt scoring can be refused with a 507 near the memory limit although it would fit (#551). Scoring did not work on Gemma 4 before this release.
  • Gemma 4 can change one token's top choice in rare chunk layouts, because its attention runs in bf16. On the M5's kernels at the default chunk size, none of seven test prompts showed it, while the single pass that 2.12.0 used showed it on two; in the M1 to M4 rehearsal on an M5 it appeared at the default chunk size. Running the attention in float32 removed it in a test, and that change is not in this release.
  • macOS 26.0 and 26.1. The MLX 0.32.2 build that installs on macOS 26 requires macOS 26.2, so the engine does not start on 26.0 or 26.1. Update macOS to 26.2 or later. macOS 14 and 15 get their own MLX builds and are not affected.

Upgrading

Run brew upgrade mtplx or pip install -U mtplx, or use the app's update check. You don't need to change any settings. On a 128 GB Mac the engine now plans with 90 GiB; --memory-limit 96G, or 96 in Settings, Memory, Memory limit, restores the old limit. pip and Homebrew now install exactly MLX 0.32.2. The hidden-tool stream guard is off; set MTPLX_STREAM_HIDDEN_TOOL_GUARD=on in the engine's environment to turn it back on. ttft_s now includes reading and parsing the request. MTPLX_PREFILL_ADMISSION_MIN_MISS_TOKENS no longer exists, because every request is now priced.

Contributors

  • First-token log-probabilities, by @jvmenen in #530
  • Prompt scoring picks its top results without a full-vocabulary pass, by @jvmenen in #532
  • Agent turns tokenize only what is new, by @jvmenen in #533
  • The session cache says why RAM could not serve a shared prompt start, by @jvmenen in #534
  • Tests no longer read your own settings, by @jvmenen in #535
  • A well-formed tool call is no longer retried as broken markup, by @jvmenen in #539
  • Blank answers at temperature 0 are no longer regenerated, by @jvmenen in #540
  • Chat requests check for Gemma 4 without building the vocabulary, by @jvmenen in #541
  • Plain chat is encoded turn by turn, by @jvmenen in #542
  • First-token log-probabilities on Gemma 4, by @jvmenen in #543
  • Prompt scoring on Gemma 4, by @jvmenen in #551
  • The reasoning repair continues the retry prompt, by @jvmenen in #552
  • MTPLX_AGENT_REWRITES=off also turns off the steering retries, by @jvmenen in #553
  • Saving a turn no longer pushes out the entry the next turn needs, by @jvmenen in #554
  • The draft head does less work during prompt processing, by @jvmenen in #555
  • Incremental SSD saves (experimental), by @jvmenen in #556
  • A stream ends at its last token, adapted from @jvmenen's #557
  • Gemma 4 sees its own tool calls in the format it is asked to use, by @jvmenen in #571
  • Dashboard axes fit their labels, by @jvmenen in #572
  • Draft heads with fused MoE experts load their experts, by @jvmenen in #574
  • Parallel tool calls use the model's own tool format, by @shiftedx in #270
  • Batched requests keep their padding out of shorter prompts, from @freddyhaddad's #522
  • Streamed replies no longer start with a blank line, by @pierre427 in #531
  • Queued SSD saves are capped, by @vahako in #536
  • The thinking guard for plain chat (experimental), by @bpmforge in #545
  • transformers up to 5.16, by @dependabot in #538

New contributors

Thanks to @boxabirds (#567 and the detailed memory report), @Lancelotbronner (#567), @nBn4u (#544), @tottemoSK (#546 and #547), @aeltawela (#525), @michel-launier (#526), @solm1121 (#528), @Deesha08 (#499), @daaku (#548), @taoeffect (#573) and @RustamKuramshin (#581) for the reports.

Full changelog: v2.12.0...v2.12.1