You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On an M5 Max, rc1 is faster than the build it replaces on almost every measure. The largest gain is tail-block reuse (#3835), which cut the next-turn wait by more than the release notes claim. Flash-Next prefill is faster, but by less than the notes say. The single-stream decode gain from #3797 did not reproduce here. Two upgrade details are missing from the notes. The depth-key rename silently changes MTP depth. Flash-Next's paged-cache block size also grows from 2,048 to 8,192 tokens.
Setup. M5 Max 128 GB, macOS 26.6.2. A is main 8a6bfb84 and B is v0.7.0rc1 35be079d. Both use mlx 0.32.2 with the same mlx-lm and mlx-vlm pins. Everything ran in one sitting, with a fresh omlx serve and an empty paged SSD cache directory per arm. MTP was on, with a single stream, temperature 0, thinking off for speed runs, and the display asleep. Models were Qwen3.8-Flash-Next-oQ4e-mtp (adaptive max depth 2 on both builds) and qwen3.8-27b-mtplx-8bit.
Measure
FN A
FN B
27B A (ceiling 3)
27B B (forced 4)
27B B (mtp_fixed_depth: 3)
Next turn of a ~10K conversation: cached tokens, TTFT
8,192, 1.57 s
10,160, 0.36 s
8,192, 3.00 s
10,160, 0.50 s
10,160, 0.50 s
Short reply (29 tokens, 2.8K prompt, warm): request to done
Read the decode row as A against B. On 2026-09-21 the same 8a6bfb84 build measured 102.6 / 75.7 on Flash-Next with the same settings and prompts, so prose ran about 14% lower in this sitting on both builds. We haven't found why yet.
Against the release notes
Tail-block reuse reproduced, with a larger effect than claimed. Re-prefill on the next turn fell from about 2,000 tokens to about 55. First token came 4x sooner on Flash-Next and 6x sooner on the 27B, against the notes' 2x (0.83 to 0.42 s).
Flash-Next prefill improved, but by less than claimed. At 16K it rose 14%, against the claimed 31.9%. At 64K it rose 20%, against 29.4%. Our runs had the paged cache on, while the notes' figures had it off. Our "before" at 64K was also already faster than theirs (1,512 against 1,326).
The single-stream decode gain from feat: faster Qwen decode with batched DFlash and improved Lightning MTP #3797 did not reproduce. Flash-Next code decode rose 3.9% and prose 0.8%, against the PR's +20.7% at batch size 1. Both of our arms capped adaptive depth at 2. If the PR number used a deeper ceiling, that may explain part of the gap.
mtp_num_draft_tokens is ignored without a warning (fix: migrate mtp_num_draft_tokens to mtp_adaptive_max_depth for backward compat #3896 adds a migration). For Flash-Next an old value of 2 falls back to the default ceiling of 3. The dense 27B gets a ceiling of at least 4 whatever the setting says. On rc1 that ceiling gave +7% code and -6.5% prose against mtp_fixed_depth: 3, which is the only way back below 4.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
On an M5 Max, rc1 is faster than the build it replaces on almost every measure. The largest gain is tail-block reuse (#3835), which cut the next-turn wait by more than the release notes claim. Flash-Next prefill is faster, but by less than the notes say. The single-stream decode gain from #3797 did not reproduce here. Two upgrade details are missing from the notes. The depth-key rename silently changes MTP depth. Flash-Next's paged-cache block size also grows from 2,048 to 8,192 tokens.
Setup. M5 Max 128 GB, macOS 26.6.2. A is
main 8a6bfb84and B isv0.7.0rc1 35be079d. Both use mlx 0.32.2 with the same mlx-lm and mlx-vlm pins. Everything ran in one sitting, with a freshomlx serveand an empty paged SSD cache directory per arm. MTP was on, with a single stream, temperature 0, thinking off for speed runs, and the display asleep. Models wereQwen3.8-Flash-Next-oQ4e-mtp(adaptive max depth 2 on both builds) andqwen3.8-27b-mtplx-8bit.mtp_fixed_depth: 3)Read the decode row as A against B. On 2026-09-21 the same
8a6bfb84build measured 102.6 / 75.7 on Flash-Next with the same settings and prompts, so prose ran about 14% lower in this sitting on both builds. We haven't found why yet.Against the release notes
content. Non-streaming/v1/messagesstill returns the reasoning as atextblock; fix(thinking): keep prompt-opened truncated thinking out of content #3893 covers that case.Two things the notes don't mention
mtp_num_draft_tokensis ignored without a warning (fix: migrate mtp_num_draft_tokens to mtp_adaptive_max_depth for backward compat #3896 adds a migration). For Flash-Next an old value of 2 falls back to the default ceiling of 3. The dense 27B gets a ceiling of at least 4 whatever the setting says. On rc1 that ceiling gave +7% code and -6.5% prose againstmtp_fixed_depth: 3, which is the only way back below 4.All reactions