Skip to content

TensorFold 0.6.1: NVFP4 checkpoints in their own math, prompts that fill together on CUDA, faster Flash Next on Macs

Choose a tag to compare

@ashhart ashhart released this 01 Oct 20:19
· 118 commits to main since this release
  • NVFP4 checkpoints in their own math. nvidia/Qwen3.8-27B-NVFP4 runs the math its checkpoint names, as vLLM
    does: 4-bit activations on Blackwell (SM 12.x), and --precision full for 16-bit activations against the same
    weights. On an RTX PRO 6000 at its 250 W limit, one stream decodes 1.4-2.0x vLLM MTP=3 and cold prompts run
    0.95-0.97x vLLM. Against an fp32 forward, top-1 agreement is 92.9% in checkpoint math (vLLM 93.0%) and 98.9% in
    full. Wider lane tiles then add 7.5-8% at 8 streams with the same tokens. RTX 40, Hopper and B200 builds are
    compiled and checked on Blackwell but have not run on those cards.
  • Waiting prompts fill together on CUDA. With --parallel, the 27B's prompts that arrive together now share one
    prefill forward: projections and MLPs run over every waiting prompt's rows, and each stream's convolution, DeltaNet
    chain and attention run on its own state, so each reply keeps its bits. On an RTX PRO 6000 at its 250 W limit, 8
    streams run 1.14-1.27x faster with NVFP4 weights and 1.26-1.36x with MLX 4-bit, and the slowest first token comes
    in 0.05-0.10 s instead of 0.4-0.8 s; on a DGX Spark, 1.11-1.19x and 1.6 s down to 0.15 s. Background and image
    prompts still fill alone.
  • More 27B tokens a second with several streams on CUDA. The stream planner prices each round on its measured
    cost, and a commit's convolution windows copy in one launch. On RTX PRO 6000 and RTX 50 cards the drafter also
    drafts as deep as the trees keep, never below the block it trained on. On an RTX PRO 6000 at its 250 W limit that
    is 1-11% more tokens a second from 1 to 8 streams, with the same tokens. A DGX Spark and other GPUs plan exactly as
    in 0.6.0.
  • Wider 4-bit lane blocks on big Blackwell cards. On SM 12.0 GPUs with 96 or more SMs, MLX 4-bit lane matmuls
    take wider blocks for 17-128 verify rows, with the same bits: 8 streams run 4-6% faster on an RTX PRO 6000 at its
    250 W limit.
  • Faster 27B decode steps on Blackwell. On a DGX Spark and on RTX 50 and RTX PRO Blackwell cards, each layer's
    4-bit projections that read the same input launch together, with tiles sized for the chip and the same bits.
    A verify step is about 4% faster on a DGX Spark and 7-15% faster on an RTX PRO 6000 Blackwell.
  • Faster 27B prompts on CUDA. The linear-attention prompt chain loads each stage's inputs one stage ahead: 1-3%
    faster prompts on an RTX PRO 6000, with the same tokens.
  • Flash Next on a Mac at long context. The sparse-attention block scores read each pooled block once for all of
    a window's rows, chained MTP drafts no longer copy the head's cache on every step, and a round's host work is
    leaner. On an 80-core M3 Ultra, one stream: code from 2% faster at 1k to 9.6% at 64k, chat level to 10.2% faster
    at 128k, with the same tokens. The sparse decode kernels are also built at load, not inside the first reply.
  • Decisions on the shared lanes. POST /v1/decisions scores a choice (labels A to Z), a score (0 to 9)
    or a yes/no from the last position's logits, without sampling a token. On Macs a decision fills on the shared
    prompt lanes while chats decode, instead of waiting for an idle engine, and keeps its prompt prefix the way a chat
    does. On an M3 Ultra, a second decision on a 6,753-token shared instruction took 2.05 s instead of 20.4 s, with
    bitwise-equal scores. A chat scored beside a decision kept its token SHA.
  • Qwen3.6-35B-A3B on Macs. The row decoder runs Qwen3.6's mixture of experts with DFlash (v1) drafts, and a
    drafter with its own draft vocabulary drafts chains sized from measured rounds.
  • Forks resume from the shared prefix. Flash Next on CUDA now keeps prompt states at message starts, so a
    conversation that forks at an earlier message resumes from the part both branches share. On a DGX Spark, a fork
    reused 755 cached tokens instead of none, and its first token came in 0.38 s instead of 0.58 s at one and two
    streams. Drafted replies equal serial ones and concurrent equal solo.
  • Shared prompt prefixes on CUDA Flash Next. A new conversation that shares a kept prefix with another, such as a
    long system prompt, now copies that prefix into a free stream instead of filling it again, and the original
    conversation keeps its own state. On a DGX Spark, 36 questions on one 40,900-token system prompt each reused it,
    and the first token came in under a second instead of 16-56 s, with the same replies.
  • Short prompts first on CUDA Flash Next. A request that arrives while a long prompt fills is admitted between
    its passes, and the prompt with the fewest rows left fills first, so a short request no longer waits for a long
    one. The long prompt's first token can come later when short ones cut in.
  • Flash Next takes images on CUDA. With --vision and --parallel of at least two on one GPU, image rows run
    through the shared prompt passes. Text replies are unchanged, and an image prompt never seeds the text prompt cache.
  • Faster Flash Next prompts on a DGX Spark. When nothing is decoding and memory allows, prompts fill in
    4,096-row pieces, so each expert weight tile serves twice the rows: 8k and 32k prompts fill about 15% faster with
    the same tokens. The hyper-connection write-back and its normalization also run as one kernel with the released
    bytes, about 2% more at 32k. On first use, each GPU checks the fused kernel against the released ones byte for
    byte and keeps the released kernels on any difference.
  • RTX cards without Docker. On an RTX 40 or 50 series card or an RTX PRO Blackwell, pip alone now installs and
    builds the CUDA kernels: torch from PyPI and NVIDIA's compiler wheels, with no root and no container. The first
    start names the compiler it found. On an RTX PRO 6000 Blackwell, the 27B with DFlash2 serves exactly: drafted
    replies equal serial ones, and concurrent equal solo.
  • Native Windows, experimental. A host layer for CUDA on Windows: one GPU a process, memory sized by Windows' own
    API, pinned buffered reads, and stack dumps on Ctrl+Break. It has not served a request on a Windows PC yet, and
    WSL2, which runs the Linux engine, is untested too.
  • Two DGX Sparks: list both RoCE devices. The direct cable shows two RoCE devices for its one port; the runbook
    now names both (NCCL_IB_HCA=rocep1s0f1,roceP2p1s0f1). With one, a 7k-token prompt read about 8% slower.
  • Metrics under vLLM's names. /metrics now mirrors its families under vLLM's names with identical values, so a
    vLLM dashboard works with a prefix swap. It also counts client disconnections and preemptions where the server
    tracks them.
  • CUDA admission on discrete GPUs. Host RAM is checked against the staging buffers a load streams, not the whole
    checkpoint, and a load that can't fit is refused before it starts. A DGX Spark's unified budget is unchanged.
  • More images per conversation. --vision-max-images N raises the image count a request may carry across its
    history above the default 4; the byte and pixel limits are unchanged.
  • Reasoning effort a template doesn't name. It now goes to the nearest level the template does name, ties going
    higher, and --reasoning-effort follows the same rule. On GLM-5.3, medium is heard as high and minimal as
    low; on Qwen3.8, high is heard as xhigh.
  • reasoning_effort: max. GLM-5.3's template names max, so max stays max there; other templates hear
    their nearest named level.
  • Independent evaluation repeats. TENSORFOLD_SEED_SALT=N mixes N into every seed taken from a prompt, so
    repeated runs of the same prompts sample independently and each stays reproducible. 0, the default, changes
    nothing.
  • mlx-lm 0.32. The Mac engine reads its cache state the same way under mlx-lm 0.31.3 and 0.32.0, with the same
    tokens on both.
  • A refused request no longer breaks the next one. Both servers answered a POST to an unknown path before reading
    its body and kept the connection open, so the next request on that connection, which behind a pooling proxy can
    come from any client, failed with 400. A refused body is now read up to the 32 MiB limit, or the connection
    closes and the reply says Connection: close.
  • Gemma 4 thought blocks stay out of replies. With thinking off, Gemma 4 can still open a thought channel, after
    a tool result or after some visible text. That block is now stripped wherever it opens, streamed or not.
  • Agentic steps on Gemma 4 reuse more of the prompt. After a tool result, Gemma 4's template adds no generation
    prompt, so no prompt state was kept near the end and each step re-read the conversation from the last turn start.
    The state now sits one token short of the end. In @gprot42's six-step agentic test, steps 2-6 reuse 78% of the
    prompt instead of 54%.
  • Flash Next on a Mac after a short request. A short request made the server overprice Flash Next's
    sparse-attention index and keep that price, so later prompts inside the window were refused. The index is now
    priced per allocated position: after a short request the fitted window is about 15% wider, and a request is
    admitted or refused the same way whatever came before it. Admission still stops short of the advertised window
    (#95 stays open).
  • Refused checkpoint captures are reported. A prompt checkpoint that memory admission refuses is now logged
    instead of dropped silently (#155).
  • GLM-5.3 EXL3 experts at any bit width. EXL3 expert layers go through the universal path at every bit width. On
    two DGX Sparks with Brandon M. Music's 4 bpw EXL3/TR3 checkpoint (re-hosted by Mia-AiLab) the tokens are the same, no timing cell is slower, and each rank
    peaks 1.7 GiB lower.
  • EXL3 grouping above 48 KiB of shared memory. The EXL3 expert grouping kernel raises the GPU's dynamic
    shared-memory limit, within the device's ceiling, before a launch that needs it.
  • Flash Next CUDA loader checks. Quantized bytes where the loader expects bf16 values are refused at load, and
    FP8 experts in the MTP drafter are read. The FP8 drafter path is checked on synthetic tensors; no checkpoint with
    FP8 drafter experts was available to run.
  • Community pull requests.
    • /v1/decisions (#127). Thanks to @mikolaj92.
    • Qwen3.6-35B-A3B on Macs with DFlash (v1) drafts (#138). Thanks to @cshintov.
    • Message-start keeps for forked conversations (#163). Thanks to @lcgutierrez.
    • Flash Next image input on CUDA (#146), adapting MiaAI-Lab's Flash Next vision patch. Thanks to @shantanugoel and
      @MiaAI-Lab.
    • Metrics under vLLM's names, and the disconnect and preemption counters (#156, #158). Thanks to @jeidbugs404 and
      @di37.
    • The nearest named reasoning effort (#161). Thanks to @EugeneClaw.
    • Short prompts first while a long one fills (#174), the loader checks and FP8 drafter experts (#176, #178), and
      seed salts (#175). Thanks to @jschmied, and to @Mirrdhyn for the shared-prefix design.
    • Configurable image history limits (#170). Thanks to @edurdias.
    • Host admission by streaming needs, EXL3 grouping above 48 KiB, and EXL3 experts at any bit width (#150, #151,
      #152). Thanks to @akol1.
    • Gemma 4 agentic prompt reuse and thinking-off thought blocks (#177, #171). Thanks to @gprot42.
    • Refused request bodies on kept connections (#182, fixing #181). Thanks to @SxMShaDoW.
    • reasoning_effort: max for GLM-5.3 (#183). Thanks to @MiaAI-Lab.
    • mlx-lm 0.32 support (#189). Thanks to @di37.

Update with tensorfold update, or pip install -U git+https://github.com/ashhart/TensorFold.git@v0.6.1.

Known limits: NVIDIA GPUs need compute capability 8.9 or newer (RTX 40, Hopper, Blackwell, the DGX Spark); RTX 30
cards are refused at startup for now. GLM-5.3 image prompts are prefilled again each turn. The discrete-GPU admission
change loads and serves on an RTX PRO 6000; its refusal path is covered by host tests.