TensorFold 0.3.6: GLM and Gemma 4 on Macs, bigger models on smaller Macs, EXL3 on NVIDIA GPUs
GLM-5.3-Flash and Gemma 4 on Macs, bigger models on smaller Macs, and EXL3 checkpoints on NVIDIA GPUs. Every drafted
reply is still byte-identical to the model's own serial decode, and resumed prompts equal fresh ones.
What's new
- GLM-5.3-Flash on Macs with 256 GB. It decodes 1.86-2.08x mlx-vlm on an M3 Ultra. Prompts process at or above
mlx-vlm from 2k to 32k tokens (1.010-1.053x), with a lower peak at long prompts (185 GiB against 213 at 32k).
Mixed-bit conversions load, and tool calls parse in both servers. Thanks to @chadhurley25075-png (#9, #39),
@jeidbugs404 (#35) and @kingjamez. - Gemma 4 26B-A4B on the lane engine, exact at every width: 9/9 drafted equal to serial, and 1 to 16 concurrent
streams equal their solo runs. Prompts process at or above mlx_lm from 2k to 64k tokens on an M3 Ultra (1.017-1.117x).
Thanks to @cshintov (#10). - Bigger models on smaller Macs.
--ple-on-ssdreads Flash Next's n-gram tables from disk: on a 128 GB budget it peaks at 85.6 GiB and decodes
at 0.91-1.03x the 256 GB run (#16).--ssd-experts GIBstreams routed experts from the checkpoint into a GPU pool of that size. Flash Next fits a
64 GB Mac (39.5 GiB peak) and GLM a 128 GB one (87.6 GiB), with the resident model's tokens. Decode runs at
0.31-0.39x resident speed for Flash Next and 0.13-0.17x for GLM. Install withpip install "tensorfold[ssd]"
(#17). Both were measured on an M3 Ultra under emulated budgets.
- EXL3 checkpoints on NVIDIA GPUs (experimental): turboderp's Qwen3.8-27B and Flash Next packs, exact on the lanes
(9/9, resume 6/6, four concurrent streams equal solo). On one DGX Spark against vLLM with MTP=3, the 27B decodes
3.56 / 1.74 / 2.50 / 1.62x and Flash Next 1.91 / 1.79 / 1.90 / 1.84x (code sampled / chat sampled / code greedy /
chat greedy). Flash Next's pack admits its full 262,144-token window on one Spark. Prompt processing is about half
the MLX checkpoints' speed for now. Thanks to @vcruz305 (#42). - Faster prompts. Flash Next sizes its prompt chunks to the memory it has: 1.20-1.38x 0.3.5.1 on an M3 Ultra,
and level with oMLX from 8k to 64k. Nemotron takes up to 8,192 tokens a chunk on M5 GPUs, 12-15% faster at 8k and
32k. The weights stay wired in memory while a server runs. tensorfold updateshows what's new when it finishes, from CHANGELOG.md, and the first run
of a new version links it.- MLX 0.32.2 or newer is required on Macs.
Fixed
- GLM-5.3-Flash on two Sparks answered "!" past about 2,000 prompt tokens: a prompt chunk's rows shared their
sparse-attention partials. EXL3 GLM prompts past 128 tokens failed too (#53, reported by @taussoe). - A reply that isn't a tool call comes back as content instead of an HTTP 500 (#51, reported by @simonmd).
- After a long prompt the server gives back 7.5-7.9 GiB at rest, clearing MLX's buffer cache only when nothing is
decoding (#44, @kingjamez).
Checks on this release
- The test suite passes: 1,778 tests on an M5 Max and 1,482 on an M3 Ultra, with no failures. The four known M1 to
M4 cases below are marked as known. - Every landed change kept drafted == serial and resumed == fresh on the machines it touches, with token hashes
equal to the tree before it wherever the output shouldn't change.
Where 0.3.6 is still short
- EXL3 prompt processing runs at about half the MLX checkpoints' speed. An FP8 prompt path is next.
- The 3x decode floor isn't met on every model: Gemma 4 is at 1.10-1.45x mlx_lm, GLM on Macs at 1.86-2.08x mlx-vlm,
and the EXL3 cells are as listed above. - Flash Next's prompt processing at 32k and 64k is 1-1.4% under oMLX on an M3 Ultra.
- On M1 to M4 with MLX 0.32.2, the 27B's prompt attention at 8,192 keys doesn't match one stock MLX call bit for bit
when a chunk splits with a short tail. 0.3.5.1 behaves the same. Drafted and serial output, and resumed and fresh
prompts, still agree. - Replies to prompts longer than one prompt chunk can differ between machines with different memory, because the
chunk size follows the memory budget. - Next in 0.3.6.1: GLM's latent cache for 256k windows (#54),
tool_choice: required(#52), explicit memory budgets
(#49, #50), the 27B at 16 concurrent streams on CUDA and the Qwen3.6-35B-A3B family (#38, #45), quantized KV and
Flash Next prefill on CUDA (#47, #40), Ternary Bonsai (#18), more quantized checkpoints, image input, and
DeepSeek-V4-Flash (#14).
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6