Skip to content

TensorFold 0.3.6.3

Choose a tag to compare

@ashhart ashhart released this 29 Sep 06:41
· 64 commits to main since this release

TensorFold 0.3.6.3

Image input on Macs and Sparks, DeepSeek-V4-Flash and Ternary Bonsai 2 on Macs, NVFP4 checkpoints on CUDA, replies
that keep decoding while long prompts prefill, and warm long conversations on 64 GB Macs. Every change keeps the
exactness rules: drafted replies equal "draft": false ones, resumed prompts equal fresh ones, and concurrent streams
equal their solo runs.

What's new

  • Images in, on Macs and Sparks. --vision lets Qwen3.8-27B read up to four images a request, as data URLs or,
    with --vision-urls, HTTPS links. Image replies keep drafting (77.6 against 37.3 tok/s serial on an M3 Ultra) and
    the first token comes as fast as mlx-vlm's. Thanks to @di37 for HTTPS-only fetching, bounded image preparation and
    redacted request logs (#64).

  • DeepSeek-V4-Flash on a 256 GB Mac, the new deepseek_v4 family, with DSpark or MTP drafts converted from
    DeepSeek's releases. On an M3 Ultra it decodes 1.7-3.1x and reads prompts 1.8-2.0x as fast as mlx-lm PR #1797's
    server. Thanks to @jeffpeng3 (#14).

  • Ternary Bonsai 2 27B on the lane engine with DFlash2 drafts: against mlx_lm running the pack's own runtime on an
    M5 Max, decode 3.7-5.8x on code and 1.7-2.2x on chat, prompts 1.4-1.7x. Thanks to @gprot42 (#18).

  • NVFP4 checkpoints on CUDA, as published: Flash Next's NVFP4 exports and NVIDIA's Qwen3.8-27B NVFP4, exact
    (84 of 84 concurrent streams equal their solo runs). On one Spark, Flash Next NVFP4 decodes 1.13-1.52x vLLM on the
    same checkpoint. Thanks to @tournierjc for the reader (#67).

  • Replies keep decoding while long prompts prefill on Macs (--decode-share, default 0.25; 0 is 0.3.6.2's
    order). On an M3 Ultra, a reply's longest pause while three 17K-token prompts arrive fell from 164 s to 7 s, and a
    request alone is unchanged. Thanks to @benwilson (#72).

  • Long conversations stay warm. On a 64 GB budget (emulated on an M3 Ultra), a conversation grew to 143K tokens
    and each turn resumed in 38-46 s, where 0.3.6.2 prefilled every turn past about 100K from the start. Thanks to
    @sanjaibalajee (#74), and to @benwilson for the report and his 64 GB measurements, now in the README (#71, #70).

  • CUDA server work from @nood-co1: stop strings on every CUDA engine and ignore_eos on the 27B (#63), 400 and
    500 answers (#61), kill -USR1 stack dumps at start and while serving (#62), and the prompt-end cache entry (#65).

  • Typed tool arguments on CUDA, thanks to @MiaAI-Lab (#75); Qwen3.6 on CUDA keeps serving past 8,192 rows
    with expandable segments, thanks to @philip-pentatonic (#78); Flash Next n-gram reads on 16 threads, thanks to
    @MovieMaker93 (#73).

  • Faster prompts on M1-M4: Flash Next's prompt kernels and expert-aligned MoE tiles (the scheduling idea of
    ml-explore/mlx#4567) make Flash Next's prompts 2-4% faster (level with oMLX at 32k and 64k on an M3 Ultra) and
    GLM-5.3's 3.4-3.8% faster, with the same bits.

  • MLX stays at 0.32.2 for now. MLX 0.32.3 changed a kernel our prompt kernels build on, so fresh installs fell
    back to slower prompts with one log line; the next release follows the new kernel. Thanks to @ecohash-co (#88).

Measured for this release

  • M3 Ultra, Qwen3.8-27B: text replies equal 0.3.6.2's token for token (18/18), image replies equal their earlier runs
    (39/39), drafted == serial, streams == solo.
  • M5 Max, Qwen3.8-27B: drafted == serial 11/11, fresh == resumed, 2 and 4 streams equal their solo runs. Decode is
    level with 0.3.6.2 within this laptop's run-to-run spread (about ±3%); one run measured code greedy at 0.977x.
  • M3 Ultra, other families against 0.3.6.2 (drafted == serial, fresh == resumed, 2 streams == solo in every one):
    Flash Next decode 1.008x and 8k prompts 1.029x; GLM-5.3 decode 0.994x and 8k prompts 1.039x; Nemotron decode
    0.980x and 8k prompts 1.006x; Bonsai is new.
  • One DGX Spark, Qwen3.8-27B: MLX 4-bit and EXL3 cells level with 0.3.6.2, --parallel 16 at 157.5 tok/s with every
    reply equal to its solo run. Qwen3.6 and Flash Next (MLX 4-bit, EXL3, NVFP4): drafted == serial, cells level.

Known gaps

  • Nemotron's chat decode on the M3 Ultra measured 2% under 0.3.6.2 (0.980x), and GLM-5.3's 0.6%; both are exact. The
    likely cause is the first drafts settling one round later (the change that makes the 27B 2.5% faster); a per-family
    setting comes next.

  • GLM-5.3 on two Sparks was not re-served for this release; its CUDA code changed only in the shared server fixes,
    which the 27B checks cover.

  • #78's crash case was not reproduced here; it rests on the contributor's run (40 of 40 on a Spark and an RTX PRO
    6000) and our Qwen3.6 non-regression check.

  • DeepSeek-V4: no M5 run on real weights yet; served checks used DSpark drafts, and MTP drafts are exact on the four
    recorded prompts. Published draft heads come in a later release; for now the recipe converts them once.