Skip to content

TensorFold 0.3.6: GLM and Gemma 4 on Macs, bigger models on smaller Macs, EXL3 on NVIDIA GPUs

Choose a tag to compare

@ashhart ashhart released this 28 Sep 09:23
· 33 commits to main since this release

GLM-5.3-Flash and Gemma 4 on Macs, bigger models on smaller Macs, and EXL3 checkpoints on NVIDIA GPUs. Every drafted
reply is still byte-identical to the model's own serial decode, and resumed prompts equal fresh ones.

What's new

  • GLM-5.3-Flash on Macs with 256 GB. It decodes 1.86-2.08x mlx-vlm on an M3 Ultra. Prompts process at or above
    mlx-vlm from 2k to 32k tokens (1.010-1.053x), with a lower peak at long prompts (185 GiB against 213 at 32k).
    Mixed-bit conversions load, and tool calls parse in both servers. Thanks to @chadhurley25075-png (#9, #39),
    @jeidbugs404 (#35) and @kingjamez.
  • Gemma 4 26B-A4B on the lane engine, exact at every width: 9/9 drafted equal to serial, and 1 to 16 concurrent
    streams equal their solo runs. Prompts process at or above mlx_lm from 2k to 64k tokens on an M3 Ultra (1.017-1.117x).
    Thanks to @cshintov (#10).
  • Bigger models on smaller Macs.
    • --ple-on-ssd reads Flash Next's n-gram tables from disk: on a 128 GB budget it peaks at 85.6 GiB and decodes
      at 0.91-1.03x the 256 GB run (#16).
    • --ssd-experts GIB streams routed experts from the checkpoint into a GPU pool of that size. Flash Next fits a
      64 GB Mac (39.5 GiB peak) and GLM a 128 GB one (87.6 GiB), with the resident model's tokens. Decode runs at
      0.31-0.39x resident speed for Flash Next and 0.13-0.17x for GLM. Install with pip install "tensorfold[ssd]"
      (#17). Both were measured on an M3 Ultra under emulated budgets.
  • EXL3 checkpoints on NVIDIA GPUs (experimental): turboderp's Qwen3.8-27B and Flash Next packs, exact on the lanes
    (9/9, resume 6/6, four concurrent streams equal solo). On one DGX Spark against vLLM with MTP=3, the 27B decodes
    3.56 / 1.74 / 2.50 / 1.62x and Flash Next 1.91 / 1.79 / 1.90 / 1.84x (code sampled / chat sampled / code greedy /
    chat greedy). Flash Next's pack admits its full 262,144-token window on one Spark. Prompt processing is about half
    the MLX checkpoints' speed for now. Thanks to @vcruz305 (#42).
  • Faster prompts. Flash Next sizes its prompt chunks to the memory it has: 1.20-1.38x 0.3.5.1 on an M3 Ultra,
    and level with oMLX from 8k to 64k. Nemotron takes up to 8,192 tokens a chunk on M5 GPUs, 12-15% faster at 8k and
    32k. The weights stay wired in memory while a server runs.
  • tensorfold update shows what's new when it finishes, from CHANGELOG.md, and the first run
    of a new version links it.
  • MLX 0.32.2 or newer is required on Macs.

Fixed

  • GLM-5.3-Flash on two Sparks answered "!" past about 2,000 prompt tokens: a prompt chunk's rows shared their
    sparse-attention partials. EXL3 GLM prompts past 128 tokens failed too (#53, reported by @taussoe).
  • A reply that isn't a tool call comes back as content instead of an HTTP 500 (#51, reported by @simonmd).
  • After a long prompt the server gives back 7.5-7.9 GiB at rest, clearing MLX's buffer cache only when nothing is
    decoding (#44, @kingjamez).

Checks on this release

  • The test suite passes: 1,778 tests on an M5 Max and 1,482 on an M3 Ultra, with no failures. The four known M1 to
    M4 cases below are marked as known.
  • Every landed change kept drafted == serial and resumed == fresh on the machines it touches, with token hashes
    equal to the tree before it wherever the output shouldn't change.

Where 0.3.6 is still short

  • EXL3 prompt processing runs at about half the MLX checkpoints' speed. An FP8 prompt path is next.
  • The 3x decode floor isn't met on every model: Gemma 4 is at 1.10-1.45x mlx_lm, GLM on Macs at 1.86-2.08x mlx-vlm,
    and the EXL3 cells are as listed above.
  • Flash Next's prompt processing at 32k and 64k is 1-1.4% under oMLX on an M3 Ultra.
  • On M1 to M4 with MLX 0.32.2, the 27B's prompt attention at 8,192 keys doesn't match one stock MLX call bit for bit
    when a chunk splits with a short tail. 0.3.5.1 behaves the same. Drafted and serial output, and resumed and fresh
    prompts, still agree.
  • Replies to prompts longer than one prompt chunk can differ between machines with different memory, because the
    chunk size follows the memory budget.
  • Next in 0.3.6.1: GLM's latent cache for 256k windows (#54), tool_choice: required (#52), explicit memory budgets
    (#49, #50), the 27B at 16 concurrent streams on CUDA and the Qwen3.6-35B-A3B family (#38, #45), quantized KV and
    Flash Next prefill on CUDA (#47, #40), Ternary Bonsai (#18), more quantized checkpoints, image input, and
    DeepSeek-V4-Flash (#14).

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6