TensorFold 0.3.6.3
TensorFold 0.3.6.3
Image input on Macs and Sparks, DeepSeek-V4-Flash and Ternary Bonsai 2 on Macs, NVFP4 checkpoints on CUDA, replies
that keep decoding while long prompts prefill, and warm long conversations on 64 GB Macs. Every change keeps the
exactness rules: drafted replies equal "draft": false ones, resumed prompts equal fresh ones, and concurrent streams
equal their solo runs.
What's new
-
Images in, on Macs and Sparks.
--visionlets Qwen3.8-27B read up to four images a request, as data URLs or,
with--vision-urls, HTTPS links. Image replies keep drafting (77.6 against 37.3 tok/s serial on an M3 Ultra) and
the first token comes as fast as mlx-vlm's. Thanks to @di37 for HTTPS-only fetching, bounded image preparation and
redacted request logs (#64). -
DeepSeek-V4-Flash on a 256 GB Mac, the new
deepseek_v4family, with DSpark or MTP drafts converted from
DeepSeek's releases. On an M3 Ultra it decodes 1.7-3.1x and reads prompts 1.8-2.0x as fast as mlx-lm PR #1797's
server. Thanks to @jeffpeng3 (#14). -
Ternary Bonsai 2 27B on the lane engine with DFlash2 drafts: against mlx_lm running the pack's own runtime on an
M5 Max, decode 3.7-5.8x on code and 1.7-2.2x on chat, prompts 1.4-1.7x. Thanks to @gprot42 (#18). -
NVFP4 checkpoints on CUDA, as published: Flash Next's NVFP4 exports and NVIDIA's Qwen3.8-27B NVFP4, exact
(84 of 84 concurrent streams equal their solo runs). On one Spark, Flash Next NVFP4 decodes 1.13-1.52x vLLM on the
same checkpoint. Thanks to @tournierjc for the reader (#67). -
Replies keep decoding while long prompts prefill on Macs (
--decode-share, default 0.25; 0 is 0.3.6.2's
order). On an M3 Ultra, a reply's longest pause while three 17K-token prompts arrive fell from 164 s to 7 s, and a
request alone is unchanged. Thanks to @benwilson (#72). -
Long conversations stay warm. On a 64 GB budget (emulated on an M3 Ultra), a conversation grew to 143K tokens
and each turn resumed in 38-46 s, where 0.3.6.2 prefilled every turn past about 100K from the start. Thanks to
@sanjaibalajee (#74), and to @benwilson for the report and his 64 GB measurements, now in the README (#71, #70). -
CUDA server work from @nood-co1: stop strings on every CUDA engine and
ignore_eoson the 27B (#63), 400 and
500 answers (#61),kill -USR1stack dumps at start and while serving (#62), and the prompt-end cache entry (#65). -
Typed tool arguments on CUDA, thanks to @MiaAI-Lab (#75); Qwen3.6 on CUDA keeps serving past 8,192 rows
with expandable segments, thanks to @philip-pentatonic (#78); Flash Next n-gram reads on 16 threads, thanks to
@MovieMaker93 (#73). -
Faster prompts on M1-M4: Flash Next's prompt kernels and expert-aligned MoE tiles (the scheduling idea of
ml-explore/mlx#4567) make Flash Next's prompts 2-4% faster (level with oMLX at 32k and 64k on an M3 Ultra) and
GLM-5.3's 3.4-3.8% faster, with the same bits. -
MLX stays at 0.32.2 for now. MLX 0.32.3 changed a kernel our prompt kernels build on, so fresh installs fell
back to slower prompts with one log line; the next release follows the new kernel. Thanks to @ecohash-co (#88).
Measured for this release
- M3 Ultra, Qwen3.8-27B: text replies equal 0.3.6.2's token for token (18/18), image replies equal their earlier runs
(39/39), drafted == serial, streams == solo. - M5 Max, Qwen3.8-27B: drafted == serial 11/11, fresh == resumed, 2 and 4 streams equal their solo runs. Decode is
level with 0.3.6.2 within this laptop's run-to-run spread (about ±3%); one run measured code greedy at 0.977x. - M3 Ultra, other families against 0.3.6.2 (drafted == serial, fresh == resumed, 2 streams == solo in every one):
Flash Next decode 1.008x and 8k prompts 1.029x; GLM-5.3 decode 0.994x and 8k prompts 1.039x; Nemotron decode
0.980x and 8k prompts 1.006x; Bonsai is new. - One DGX Spark, Qwen3.8-27B: MLX 4-bit and EXL3 cells level with 0.3.6.2,
--parallel 16at 157.5 tok/s with every
reply equal to its solo run. Qwen3.6 and Flash Next (MLX 4-bit, EXL3, NVFP4): drafted == serial, cells level.
Known gaps
-
Nemotron's chat decode on the M3 Ultra measured 2% under 0.3.6.2 (0.980x), and GLM-5.3's 0.6%; both are exact. The
likely cause is the first drafts settling one round later (the change that makes the 27B 2.5% faster); a per-family
setting comes next. -
GLM-5.3 on two Sparks was not re-served for this release; its CUDA code changed only in the shared server fixes,
which the 27B checks cover. -
#78's crash case was not reproduced here; it rests on the contributor's run (40 of 40 on a Spark and an RTX PRO
6000) and our Qwen3.6 non-regression check. -
DeepSeek-V4: no M5 run on real weights yet; served checks used DSpark drafts, and MTP drafts are exact on the four
recorded prompts. Published draft heads come in a later release; for now the recipe converts them once.