Skip to content

Releases: ValerioDolci/ninfer-tp2

v0.4.7

Choose a tag to compare

@ValerioDolci ValerioDolci released this 05 Oct 17:00

Fifteenth release, without new binaries: the C++ tree is the same as v0.4.6, so the v0.4.6 tarball runs it. Changes since v0.4.6 (47d47864 → v0.4.7):

  • MTP layer in NVFP4, as an opt-in conversion recipe. _mtp_nvfp4_layer in tools/convert/official_recipes.py (with tests) stores the MTP layer's MLP gate/up, MLP down and attention output as NVFP4 (nvfp4_mse from the source's BF16 MTP weights, 16-bit activations); at tp 2 their halves run the NVFP4 linear routes the target already uses. The QUASAR model card recipe gains configure_mtp_nvfp4 (--recipe quasar_recipe.py:configure_mtp_nvfp4); it reproduces the measured artifact object for object (1,524/1,524). Against the published artifact on 2× RTX 5070 Ti (eco clocks, MTP3, --lm-head-draft, 262,144 tokens, T=0): the MTP round is 1.9 % shorter at C=1 and 1.4 % at C=4; outputs, target perplexity and GSM8K (500 problems, thinking on: 483/500 on both, the same answer on every problem, McNemar p = 1.000) are identical; decode is 1.5 % faster on average at C=1 (2.2–2.9 % on code, math and English prose, 0.6 % on Italian, unchanged on agent prompts of 55–171K tokens, where acceptance drops by 1.7 points). Weights −142.5 MiB, +68 MiB of free VRAM on rank 0. The production profile of this fork now runs this variant (qwen3_8_27b_quasar_nvfp4_mtpnv.ninfer, SHA-256 3363b1ab64c5693c8dcb0cb283b01bcb1a19de946846799e53bca7ac8f2533ac); the artifact on the Hub is unchanged.
  • Gate: a paired perplexity test for commits that change bits on purpose (GATE_PPL_PAIRED=1, tools/tp2/ppl_paired.py, standard library only). An identical table still passes as before; a different one is compared window by window with the reference report.json (mean ΔNLL per token, standard error across windows, paired t, 95 % interval of the perplexity change, exact sign test): PASS if not worse at p > 0.05 and |Δ| < 0.5 %, WARN otherwise, FAIL if the windows differ. --record also stores ppl-*.report.json. Documented in tools/tp2/README.md with the GSM8K/McNemar rule for bit changes outside the gate.
  • Gate reference recorded again on the MTP-NVFP4 artifact (GATE_ARTIFACT default …_mtpnv): golden 3/3, perplexity tables identical to every digit (the per-window reports differ only in timings and the artifact path: perplexity measures the target, which this artifact does not change), MTP3 greedy 60/60 texts identical at 16.693 ms/round (was 17.000) with 73.70 % acceptance (73.62 %), DFlash2 unchanged (10/10, 18.204 ms/round).
  • Two other draft-head reductions were measured and not adopted: a 65,536-row MTP proposal head (acceptance −3.5 points on average, −8.9 on long agent prompts, and its tp 2 half has no linear route) and the DFlash2 drafter's down projection in NVFP4 (+0.6 % decode, below the 1 % bar; it also does not run with the whole drafter on rank 0).

Verified: conversion tests; gate of the candidate artifact on the v0.4.6 runtime (shards PASS, golden 3/3, perplexity identical to every digit, MTP3 greedy 60/60 texts identical, 16.694 vs 17.000 ms/round); paired-perplexity check end to end on a build that changes bits; production smokes after the switch (text, thinking, image).

Not changed: the runtime (C++/CUDA), defaults, the artifact format, the published artifacts.

v0.4.6 — request log labels and telemetry, speculation detail option, residual rounding independent of tile occupancy

Choose a tag to compare

@ValerioDolci ValerioDolci released this 05 Oct 13:39

Fourteenth packaged release. Changes since v0.4.5 (ff6b336f → v0.4.6):

  • Request log (--request-log-jsonl, schema 21 → 27):
    • request.client carries the optional X-Ninfer-Client header (1–64 characters of A-Z a-z 0-9 . _ : / @ + -; a malformed or repeated header is a 400 invalid_client_label), so benchmarks and smokes can be told apart from real traffic. The fork's own tools (gate client, corpus, concurrency, TTFT, agentic A/B, smokes) now send <kind>/<tool> labels.
    • speculative.accepted_length_histogram: how many rounds accepted exactly i draft tokens.
    • model_thinking_tokens is counted for every response that starts in thinking; it was only counted when a thinking budget was set, so without one it read 0.
    • throughput.gpus: per board, SM and memory clocks, power, energy over the interval, temperature and the clock-throttle reasons, read through NVML (loaded at runtime with dlopen, matched by UUID) on the statistics thread: 4.2 ms per sample of two boards every 5 s, nothing on the request path.
    • --log-speculation-detail (off by default; --log-speculation-detail-max-steps N, default 4096): adds speculative.accepted_lengths, one character per generation step in order (a hex digit = draft tokens accepted in that round, - = a step without a draft; accepted_lengths_truncated past the cap), and result.prompt_ngram_overlap, the share of the output's 4-token windows that also occur in the prompt (thinking and tool calls included; 0.016 ms at 8k prompt tokens, about 1 ms at 262k, on the finishing request only). Off, the decode round is unchanged (16.996 vs 16.998 ms); the public API gains PreparedPrompt::token_ids().
  • Fused residual rounding no longer depends on tile occupancy (NVFP4 A4 and FP8 A8 linear_add, the prefill routes). The FullTokens kernel instance, chosen only when the launch width is an exact multiple of the tile, let nvcc contract the scaled accumulator and the residual into one FMA while every other launch (tail tiles, the 1024-token TMA chunks) rounded twice, so a column's result depended on the batch width. Reported and fixed by MirkoCovizzi (87f0689805, 37a1ed49d4); ported here keeping, per format, the rounding the long chunks already used — NVFP4 explicit MUL then ADD, FP8 FMA — so the fix changes nothing that the gate reference measured: perplexity identical to every digit at tp 2 and tp 1, golden 3/3, MTP3 greedy 59/60 texts (one prompt whose prefill segment was tile-aligned), DFlash2 10/10, GSM8K 1,319 problems 96.59 % vs 96.44 % (McNemar p = 0.50), speed unchanged. A new CTest (ninfer_linear_add_residual_rounding_test) checks a residual column keeps its bits at every batch width (9 NVFP4 and 14–21 FP8 widths differed before). Mirko's literal port (FMA everywhere) was measured and not adopted: +0.10 % perplexity at 64k, 14/60 texts.
  • Gate reference re-recorded on this tree (greedy 17.000 ms/round; the DFlash2 stage now defaults to the NVFP4-drafter artifact _df2nv: 18.200 ms/round, 62.66 % acceptance); ninfer_resource_manager_test compiles again and covers the cancelled-request path (it had been excluded from the ctest stage since v0.4.4).
  • Documentation: the README summarizes the 2026-10-05 changes; docs/serving.md, docs/maintainer/logging.md and docs/weight-conversion.md cover the new options, the log schema and the NVFP4 drafter gate/up conversion.

Verified: ctest 148/148 on the rounding tree; request-log, transport, options, public-API and frontend tests; the two-device MTP engine test for the recorded accepted lengths; behavior gate on main with and without --log-speculation-detail (golden 3/3, perplexity identical to every digit, MTP3 greedy 60/60 texts identical at 16.996–16.999 ms/round, DFlash2 10/10); HTTP smokes on the chat, Anthropic and Responses APIs with the detail on. The release binary's gate is recorded in the production dossier.

Not changed: the decode (A16) kernels, defaults (every new option is off), the artifact format.

Tarball: same layout as v0.4.5 (stripped ninfer, ninfer-serve, ninfer-perplexity, CUDA 13.2 runtime, serve-tp2.sh, SHA256SUMS; Ubuntu 26.04, sm_120a).

v0.4.5 — token embedding in host memory, Anthropic thinking.display, NVFP4 DFlash2 drafter MLP

Choose a tag to compare

@ValerioDolci ValerioDolci released this 05 Oct 08:55

Thirteenth packaged release. Changes since v0.4.4 (ae667d61 → v0.4.5):

  • --embedding-host: the token_embedding table (FP8 rows, 1.18 GiB) is kept in one mapped page-locked host buffer and read by both boards over PCIe instead of being replicated in each board's VRAM. The gather kernels are unchanged, so results are identical bit for bit (eight full gates, with and without the option, all identical to the reference). Frees 1.18 GiB per board: with the production flags the free VRAM goes from 920 MiB to 2.08 GiB on qwen27b (MTP3, 262,144) and from 208 MiB to 1.38 GiB on qwen27b-code (DFlash2, 229,376). Cost: prefill −0.9 % at 8k and −0.5 % at 115k tokens (every board reads 5 KiB per prompt token over PCIe); decode rounds unchanged (MTP3 −0.02 %, DFlash2 +0.13 %, 3 runs each, 2.1 GHz cap). What it buys: the DFlash2 profile reaches the full 262,144 context (a 244,725-token request served), MTP3 at four concurrent requests takes --kv-capacity 327680 int8 (ceiling ≈360k), and at --tp 1 the 27B artifact fits one 16 GB board for text (≈22–26k context, estimated). Default off. Not exercised: vision prompts, WSL2.
  • Anthropic thinking.display: thinking: {type: adaptive, display: omitted} is accepted (Claude Code 2.1.289 sends it on every request together with context_management; v0.4.4 answered 400 and the client stopped at its first turn). Ported from Wallawalla47/ninfer-custom ca1f50ed. Verified with the schema tests, an HTTP smoke (11/11) and a real two-turn Claude Code session with a tool call.
  • DFlash2 drafter MLP gate/up as NVFP4 (ported from ninfer-custom 1f2f5e97, on the rank-split drafter). Needs a reconverted DFlash2 artifact (5 of 1,590 objects change, −425 MiB; 4 minutes of CPU with the conversion tools); with the existing artifact nothing changes. Measured at C=1 against v0.4.4: decode +2.1 % (code), +2.4 % (math), +1.6 % (prose), +1.5 % (agent 55k), −0.5 % (agent 128k); acceptance within ±1.2 points; round −1.3…−1.7 %; free VRAM on rank 0 at 229,376 from 208 to 420 MiB. At C=4 the prose round is +0.6 % (a 17–32-token route the fork never tuned for 70-SM boards).
  • Artifact loader: the direct-I/O staging slots are aligned to the payload alignment themselves (a small page-locked allocation is not always 4,096-byte aligned; the 64 MiB production slots were).

Verified: behavior gate PASS on the ported branch and on --embedding-host on/off (shards, golden 3/3, perplexity 64k/4k identical to every digit, MTP3 greedy 60/60 texts identical at 16.99 ms/round, DFlash2 10/10 identical on the existing artifact); the release binary's gate is recorded in the production dossier.

Not changed: kernels of the dense model, defaults (both --embedding-host and the NVFP4 drafter are opt-in: the latter by artifact), the gate reference.

Tarball: same layout as v0.4.4 (stripped ninfer, ninfer-serve, ninfer-perplexity, CUDA 13.2 runtime, serve-tp2.sh, SHA256SUMS; Ubuntu 26.04, sm_120a).

v0.4.4 — a cancelled or edited turn keeps the conversation cache

Choose a tag to compare

@ValerioDolci ValerioDolci released this 05 Oct 07:07

Twelfth packaged release. Changes since v0.4.3 (7f500254 → v0.4.4):

  • A cancelled turn no longer costs the conversation its cache (issue #3). A same-session follow-up consumes the previous turn's continuation, and a cancelled request used to be discarded with it; on top of that the follow-up's own turn-closure checkpoint replaces the previous one, so after a cancellation (or an edited last message) the next prompt branched off a point with no checkpoint left and rebuilt the whole conversation. Two changes:

    • a cancelled request is now catalogued at its committed frontier like a finished turn (Program::finish_cancelled; the discard remains the fallback), so sending the same prompt again resumes from its closure or endpoint;
    • --turn-anchors N (default 1, 0 disables): the server plants a private long-anchor opportunity at the message boundary before the final message (with N > 1, before the preceding messages too), so the follow-up's prefill keeps a checkpoint exactly where a cancelled or edited turn branches. It uses the existing long-anchor machinery within --max-long-anchors-per-continuation; one Device state slot per conversation.

    Measured on a 34k-token conversation, two RTX 5070 Ti, MTP3, thinking on: cancel during decode then a different follow-up 6.85 s → 0.09 s of TTFT (34,148 of 34,178 tokens from the cache); cancel during prefill 6.86 s → 0.07 s; an edited last message without any cancellation 6.85 s → 0.08 s; the cancelled prompt sent again 0.02–0.05 s. Thinking off behaves the same.

  • Speculative sampling (upstream issue Neroued#349): a draft whose proposal-probability lookup misses (q = 0) is rejected instead of accepted unconditionally; the residual resample then draws from the unmodified target distribution. The contract makes the miss practically impossible (~3e-9 per draft); greedy decoding is unaffected.

  • Upstream fixes: explicit KV publication streams (064965c7, the Neroued#320 Xid 79 class), and a synchronization of both devices when a request fails to start (from 75a89050; the abort half was already covered in v0.2.1).

  • README: the WSL2 note carries the v0.4.2 report from issue #4 (MTP3 at 112–114 tok/s, the full 262,144 context with --kv-dtype nvfp4, pipelined mailbox starting cleanly).

Verified: behavior gate PASS on this tree (shards 28, golden 3/3 identical, perplexity 64k 4.133481 and 4k 4.385905 identical to every digit, MTP3 greedy 60/60 texts identical, DFlash2 10/10 identical; ms/round not comparable to the reference this time, the boards ran at free clocks: 14.00 ms against the 17.14 ms recorded at the 2.1 GHz cap, the known −18 %). The issue-#3 scenarios above were measured with the fork's reproduction script before and after each change.

Not changed: kernels, defaults other than --turn-anchors 1, the artifact format.

Tarball: same layout as v0.4.3 (stripped ninfer, ninfer-serve, ninfer-perplexity, CUDA 13.2 runtime, serve-tp2.sh, SHA256SUMS; Ubuntu 26.04, sm_120a).

v0.4.3 — faster tp 2 down projection on 70-SM boards

Choose a tag to compare

@ValerioDolci ValerioDolci released this 03 Oct 09:06

Eleventh packaged release. Changes since v0.4.2 (4d951afa → 7f500254):

  • Faster tp 2 down projection on 70-SM boards (7f500254). The two-device down half [5120,8704] (NVFP4, A16) runs four rows per warp instead of two at T=3..5 — on both the plain linear() route (rank 1) and the fused residual-add route (rank 0) — and eight K-warps instead of four at T=6..8. Measured on an RTX 5070 Ti with ninfer_linear_bench (medians of 3×60): T=3/4 46.8 → 44.7 µs (−4.5 %), T=5 53.0 → 46.8 µs (−11.6 %), T=6..8 53.0 → 50.9 µs (−3.9 %). On the MTP3 decode round at tp 2 (60 greedy prompts, two runs): 17.137 → 17.017 ms per round (−0.7 %), generated texts 60/60 identical; DFlash2 10/10 identical.

Verified: behavior gate PASS on this tree (shards 28, ctest 148/148, golden 3/3, perplexity 64k/4k identical to every digit, MTP3 greedy 60/60 texts identical, DFlash2 10/10 identical).

Not changed: the other NVFP4 shapes. The same study found the large GEMVs (gate-up, GDN in-projection, QKV) already at 86-95 % of the board's measured DRAM read ceiling (693 GB/s with the core clock capped at 2.1 GHz, 855 GB/s with free clocks), and programmatic dependent launch on the A16 routes within noise (−0.4 %).

Tarball: same layout as v0.4.2 (stripped ninfer, ninfer-serve, ninfer-perplexity, CUDA 13.2 runtime, serve-tp2.sh, SHA256SUMS; Ubuntu 26.04, sm_120a).

v0.4.2 — NVFP4 KV cache at tp 2

Choose a tag to compare

@ValerioDolci ValerioDolci released this 03 Oct 07:20

Tenth packaged release. Changes since v0.4.1 (2480b707 → 4d951afa):

  • NVFP4 KV cache at tp 2 (be1960ce, proposed and verified on 2× RTX 5080 by @glfenix, issue #2). --kv-dtype nvfp4 is accepted with --tp 2: the NVFP4 causal-attention kernels are instantiated for the per-board head split [D=256, Hq=12, Hkv=2] (grouped prefill, parallel decode, tiled MMA) and the tp 2 gates accept Nvfp4Group16. No change on the bf16/int8 paths. The default stays int8.
  • Two-device head-local parity test also for NVFP4 (4d951afa): ninfer_attention_headlocal_test now compares the [256,24,4] geometry on one board with the [256,12,2] halves on two ranks for bf16, int8 and nvfp4.

Measured on two RTX 5070 Ti 16 GB (QUASAR-QAT NVFP4 27B, no P2P, clocks ~2.1 GHz):

  • Quality, nvfp4 vs int8: perplexity 65536/32768 4.1229 vs 4.1335, 4096/2048 4.3987 vs 4.3859; GSM8K-500 (C=1, thinking) 0.974 vs 0.972.
  • Memory with the production options at 262,144 context: 2.89 GiB free per board vs 920 MiB with int8 (~2 GiB saved). --max-context above 262,144 is rejected on this artifact by its position capacity with either dtype; the saving can go to --kv-capacity (393,216 and 524,288 start with nvfp4 at 262,144 context; int8 cannot exceed 262,144).
  • Speed: short prompts on par. From 55k to 210k prompt tokens the round costs +4…8 % with MTP3 and +4…17 % with DFlash2 (whose acceptance also drops 10–15 points); TTFT +1…2 %. Hence int8 remains the default on boards where the int8 cache fits.

Verified: behavior gate PASS on this tree (shards 28, ctest 148/148, golden 3/3, perplexity 64k/4k identical to every digit, MTP3 greedy 60/60 texts identical at +0.00 % ms/round, DFlash2 10/10 identical); ninfer_softmax_attention_test --kv-dtype nvfp4 PASS including the 12/2 cases.

Tarball: same layout as v0.4.1 (stripped ninfer, ninfer-serve, ninfer-perplexity, CUDA 13.2 runtime, serve-tp2.sh, SHA256SUMS; Ubuntu 26.04, sm_120a).

v0.4.1 — wide mailbox slots for 2-4 concurrent requests, DFlash2 selector on rank 1

Choose a tag to compare

@ValerioDolci ValerioDolci released this 01 Oct 06:09

Ninth packaged release. Changes since v0.4.0 (60c47e45 → 2480b707):

  • Wide mailbox slots for batched rounds (62839967). The mailbox slot carried one request's widest exchange, so at 2-4 concurrent requests every all-reduce of a round (128 in the verification, 10 in the split DFlash2 drafter) fell back to staged copies inside the graph. With the pipelined kernel the mailbox now also has slots for a full batch, in a pinned slab of their own (one-request exchanges keep the small slab and stay exactly as fast). Cost: 1.3 MiB of pinned host memory, no VRAM. NINFER_TP_MAILBOX_SLOT=request restores the v0.4.0 behavior on the same binary.
  • DFlash2 candidate selector on rank 1 (fdf6c33a). The selector (257 MB) moves off GPU 0, which bounds the KV capacity. It stays on rank 0 with vision on GPU 1, with NINFER_TP_DRAFT_HEAD=primary, or with NINFER_TP_SELECTOR=primary (full DFlash2 speed when the extra context is not needed).
  • The behavior gate's reference now lives in the tree (tools/tp2/reference/, 0a590655, 2480b707).

Measured on two RTX 5070 Ti (QUASAR NVFP4, production options: KV int8, 196,608 context, 4 slots, vision on GPU 0; alternated A/B against v0.4.0):

  • MTP3: 2 concurrent requests 259.8 → 293.1 tok/s (+12.8 %), 4 concurrent 439.1 → 481.0 (+9.5 %); 1 request unchanged.
  • DFlash2: 2 concurrent 295.8 → 329.5 tok/s (+11.4 %), 4 concurrent 482.7 → 542.8 (+12.5 %); 1 request unchanged.
  • DFlash2 with the selector on rank 1: max context 233.3k → 248.5k tokens (a 237,631-token request served), at +0.5 % time per round with 1 request and +1 % at 2-4.
  • Generated texts at 1 request identical to v0.4.0 (15/15, MTP3 and DFlash2); drafts bit-identical with the selector on either rank.

Verified: behavior gate PASS (ctest 148/148, golden, perplexity identical to every digit, MTP3 greedy 60/60 identical, DFlash2 10/10 identical); GSM8K 200 with DFlash2 0.970, 200/200 answers identical to v0.4.0.

Tarball: stripped ninfer, ninfer-serve, ninfer-perplexity, the CUDA runtime they link (13.2), serve-tp2.sh and SHA256SUMS; links the distribution's FFmpeg 8 and libcurl (Ubuntu 26.04, CUDA 13.2-capable driver ≥ 595). Elsewhere build from source.

v0.4.0 — DFlash2 drafter split across the two GPUs

Choose a tag to compare

@ValerioDolci ValerioDolci released this 30 Sep 21:46

Eighth packaged release. Changes since v0.3.0 (169514ea → 60c47e45):

  • DFlash2 drafter split across the two GPUs (8f1540c3, 1d32d3e6, 60c47e45). The optimized proposal head splits by vocabulary with an exact top-16 merge; the drafter layers split like the target (attention 16/4 heads per rank, MLP halves, row-parallel sums with an identical residual on both ranks), with rank 1 keeping its own drafter state. Until now the whole drafter ran on GPU 0 while GPU 1 waited (4.1 ms, 20 % of a round). NINFER_TP_DRAFTER=primary restores the single-GPU drafter on the same binary.
  • The two-GPU layer now lives in tp2/ directories (19 commits, bit-identical): our code inside upstream files went from +12,867 to +6,690 lines, and replaying the last upstream sync gives 14 content conflicts instead of 26. One per-GPU tuning table (src/core/tp2/device_tuning.h), a behavior gate (tools/tp2/gate.sh), a merge guide with "Adding a new GPU" and "Adding a new model to tp2" (docs/maintainer/upstream-merge.md).

Measured on two RTX 5070 Ti (QUASAR NVFP4 + z-lab DFlash2, --spec dflash2 --draft-tokens 7 --lm-head-draft, KV int8, alternated A/B against v0.3.0):

  • DFlash2 decode, 1 request: code 261.3 → 282.4 tok/s (+8.1 %), maths 298.1 → 322.1 (+8.0 %), prose 119.3 → 129.7 (+8.8 %), 64k context +7.0 %; 4 concurrent requests +4.5 %; time per round 19.97 → 18.44 ms; generated texts identical, acceptance within ±0.2 pp.
  • DFlash2 max context with vision on GPU 0, 4 slots and 8 device-state slots: 157.5k → 233.3k tokens (a 214k-token request served).
  • LocalMaxxing reasoning-v1, uncapped clocks, 4 runs each: 290 → 328 tok/s.
  • MTP3 (the default mode): unchanged, bit-identical to v0.3.0.

Verified: behavior gate PASS (ctest 148/148, golden, perplexity identical to every digit, MTP3 greedy 60/60 identical, DFlash2 10/10 identical); GSM8K 200 with DFlash2 0.970 on both builds with 200/200 identical answers; the refactor alone: GSM8K 500 C=1 500/500 identical to v0.3.0.

Tarball: stripped ninfer, ninfer-serve, ninfer-perplexity, the CUDA runtime they link (13.2), serve-tp2.sh and SHA256SUMS; links the distribution's FFmpeg 8 and libcurl (Ubuntu 26.04, CUDA 13.2-capable driver ≥ 595). Elsewhere build from source.

v0.3.0 — upstream d44ab584 + INT8 attention tuned for the RTX 5070 Ti

Choose a tag to compare

@ValerioDolci ValerioDolci released this 30 Sep 17:26

Seventh packaged release. Changes since v0.2.3 (629b2a34 → 169514ea):

  • Rebased on upstream Neroued/ninfer d44ab584 (merge 609e5c92): upstream's reorganized causal attention (one directory per KV format), FP8 linear tuning and native FP8/NVFP4 conversions. The two-GPU layer was re-ported onto it: 12/2 head-local attention for BF16 and INT8 KV (cbd0f83d), kernel attributes set on every device (0458a241), FP8 row-parallel halves re-derived from upstream's retuned parents (33cda4b3), and a fix for FP8 split-K scratch sizing on the halves (6ae0586e).
  • INT8 attention budgets split-KV on the device's own SM count (4b9c098e, 21ad2f4a): the plan sized its waves for 170 SMs (RTX 5090); on a 70-SM RTX 5070 Ti that was 2.4 waves. Now queried once per device; unchanged on a 5090. Kernel geomean −10 to −16 % at 12/2, 24/4 and 16/2 on the 5070 Ti.
  • Short-context split rule (169514ea): a row takes as many splits as it has 32-key tiles, up to one wave of the device, balanced beyond it. Short rows −15 to −40 % at kernel level (T=4 verify at 1k keys −40 %).
  • Build with nvcc ≥ 13.2 (this tarball: 13.2.86): enables upstream's native conversions and is 0.5-2 % faster than 13.1 on this code; the bundled CUDA runtime needs a CUDA 13.2-capable driver (≥ 595).

Measured on two RTX 5070 Ti (QUASAR NVFP4 weights, INT8 KV, MTP3 + --lm-head-draft, paired A/B against v0.2.3 in alternated order):

  • time per MTP round: −0.3 % at 0k, −2.8 % at 16k, −3.5 % at 64k, −4.2 % at 120k; cold prefill +0.2…+0.7 %; 60 short prompts 185.1 vs 182.6 tok/s (+1.4 %);
  • LocalMaxxing reasoning-v1, uncapped clocks, 3 runs each: MTP3 222.9 vs 218.9 tok/s, DFlash2 K=7 302.2 vs 291.0 (single-run spread 2.5-5.7 %).

Verified: perplexity at --tp 2 bit-identical to v0.2.3 (4.133481 at 64k windows, 4.385905 at 4k); GSM8K 500 C=1 0.976 (v0.2.3: 0.972, McNemar p = 1.00) and C=4 0.980; MMLU-Pro 308 0.776 and IFEval 200 0.900 on the merge (within noise of the previous build); GPU test suite 146 pass / 1 skip / 7 fail (the 7 are tp1 27B out-of-memory cases on a 16 GB card), two-GPU real tests 10/10, tp1 golden identical to upstream on the BF16/FP8 cases.

Tarball: stripped ninfer, ninfer-serve, ninfer-perplexity, the CUDA runtime they link, serve-tp2.sh and SHA256SUMS; links the distribution's FFmpeg 8 and libcurl (Ubuntu 26.04, CUDA 13.2-capable driver). Elsewhere build from source.

v0.2.3 — split proposal head, faster A16 SwiGLU/down on 70-SM GPUs (bit-identical); matches the RTX 5090 on MTP3 structured output

Choose a tag to compare

@ValerioDolci ValerioDolci released this 28 Sep 07:57

Sixth packaged release. Changes since v0.2.2 (58fcdd1f → 629b2a34):

  • Optimized proposal head split across the two ranks (491e2a80): with --lm-head-draft, each rank now projects half of the Q4 draft head (65,536 rows) and takes the argmax of its half; the two candidates are exchanged in one 16-byte message per column on the same all-reduce and transport as the rest of the round, and rank 0 keeps rank 1's candidate only if strictly larger, so ties resolve to the lower row exactly as the whole-head argmax did. Drafts, MTP acceptance and output text are identical; −4.0 / −3.8 / −3.5 % per round at 0 / 16K / 64K (118 → 123 tok/s at 0K, one request); 170 MiB move from rank 0 to rank 1. NINFER_TP_DRAFT_HEAD=primary restores the whole head on rank 0. DFlash2 keeps the whole head on rank 0 as before.
  • Faster NVFP4 A16 SwiGLU and down projections at small T on 70-SM GPUs (91582814): at T=3..8 the sliced-K kernel ran three CTAs per SM (78 registers, 25.6 KiB of shared memory holding eight activation rows even for four tokens); it now loads four rows with 64 registers at T≤4 and takes two row tiles per CTA at T=5..16, and the SIMT down projection runs six CTAs per SM at T=3..5. Same K slices, same MMA sequence, same reduction: 356/356 measured points bit-identical. SwiGLU [17408,5120] at T=4: 85.0 → 78.2 µs (its A4 route: 78.9); the full [34816,5120] parent 166 → 150 µs, so this also helps --tp 1 on GPUs with few SMs. −3.3 / −2.9 % per round at 0 / 16K.
  • Uncapped-clock numbers (629b2a34, README and docs/performance/two-gpu.md): on two RTX 5070 Ti at about 2.9 / 2.8 GHz under load, with the official weights the pair matches the published RTX 5090 figures on MTP3 structured output (220.0 vs 219.8 tok/s) and on plain decode at 7.7K (70.9 vs 71.2), and reaches 98 % on the MTP3 corpus at C=1 (157.1 vs 161.1); prefill stays at 59 % (4,956 vs 8,340). With the QUASAR-QAT artifact: structured output 254.5 tok/s (116 %), plain decode 118 %, prefill 72 %. Against llama.cpp on the same boards (official weights, same tests as 2026-09-24): 2.1× on a cold 16K prompt, 1.9× decode at 184K context, a 16-turn agent session 44 % shorter.

Verified: GSM8K 500 questions at C=1 with --lm-head-draft: 0.972 with 500/500 answers identical to v0.2.2; two-GPU test suite and decode_graph 11/11 on this build; smoke (every reasoning_effort, chat, tool calls); three parallel requests through the production queue at C=4.

Tarball: stripped ninfer, ninfer-serve, ninfer-perplexity, the CUDA runtime they link, serve-tp2.sh and SHA256SUMS; links the distribution's FFmpeg 8 and libcurl (Ubuntu 26.04, CUDA 13.1-capable driver). Elsewhere build from source.