Skip to content

Releases: poisonxa16/pxa

PXA v2026.10.2: big context on one card, Gemma 4 MTP

Choose a tag to compare

@poisonxa16 poisonxa16 released this 30 Sep 16:37

PXA Network

Discord Ko-fi Free model

Install in one line: curl -fsSL https://raw.githubusercontent.com/poisonxa16/pxa/main/install.sh | bash · Docker: docker pull ghcr.io/poisonxa16/pxa:v2026.10.2

A fix for large-context one-card decode, Gemma 4 speculation that works, and a lighter Gemma KV cache. Same models, no re-quant. Upgrade recommended.

Fixed: a large context could silently push weights out to host RAM

On a single 16 GB card, an f16 KV cache at a large context (for example -c 65536 on the Qwen3.8-27B one-card file) tipped the weights from "all on the card" to "part read from host RAM over PCIe on every step". Decode fell from about 24 to 1.7 t/s and prefill from 50 to 14 t/s. It was not a kernel bug, it was a placement cliff.

  • With no -ctk / -ctv given, the engine now picks the fastest KV type that keeps the weights on the card: f16 if it fits, else q8_0, else q4_0, and logs the arithmetic on the PXA_REGISTRY line.
  • One P100, Qwen3.8-27B one-card file, no KV flags: -c 8192 picks f16 (24.7 t/s), -c 32768 picks q8_0 (24.4 t/s), -c 65535 picks q4_0 (24.2 t/s). Everything stays resident. One V100 at -c 65535 picks q4_0 and decodes at about 33 to 35 t/s.
  • If a placement still has to spill, the server now says so in the log instead of just getting slow.
  • Anything you pass yourself (-ctk, -ctv) is used as given.

Gemma 4 26B-A4B: speculative decoding works, and the KV cache is much smaller

  • MTP with the upstream drafter. The Gemma 4 assistant (drafter) file now loads as it is published: no conversion. Pass it with -md and the engine arms --spec-type mtp:n_max=1 by itself. On one V100 (QAT q4_0, 16k context, greedy), long prose goes from 117 to about 135 t/s and code to about 134 t/s; with --spec-type mtp:n_max=2 prose reaches about 150 t/s (149 to 153 across runs). UD-Q4_K_XL: 118 to about 136 t/s prose, about 120 code.
  • Sliding-window KV is the default. Gemma 4 keeps only the window it needs for its local layers: the KV cache at 16k context drops from 3.5 GB to 0.77 GB. On a 16 GB card the old build could not even start the server at -c 8192 and -c 16384 with the QAT file; this one boots at 16k and recalls a needle placed in a 14k-token prompt. PXA_GEMMA4_ISWA=0 restores the old layout.
  • V100 prefill on long Gemma prompts is about 3% faster (pp4096 1986 to 2044 on the QAT file, 1997 to 2055 on UD-Q4_K_XL). pp512 and plain decode (about 115 to 120 t/s) are unchanged.

Limits, plainly

  • Non-English text is where the drafter agrees least. On a Finnish answer only about 40% of drafted tokens are accepted, so MTP does not help there. A gate stops drafting when acceptance drops under 0.5 and keeps decode near plain speed, but "near" means about 10% below it on Finnish (about 104 vs 116 t/s), and short English chat answers are a few percent below plain as well (111 vs 118). Long prose and code are where MTP pays. PXA_GEMMA4_MTP_GATE=0 turns the gate off, PXA_GEMMA4_MTP_AUTO=0 stops -md from arming MTP by itself.
  • A Gemma 4 prompt of about 16k tokens still runs out of memory on a 16 GB card at the default batch size. 14k tokens is what we verified on one V100.
  • The drafter file we tested is the q4_0 requant of the upstream assistant.

One P40 item (unmeasured)

A single sm_61 card with 20 GiB or more (a 24 GB P40) was being treated like an 11 GB 1080 Ti: small batch, speculation declined. It now gets the 16 GB-class batch (-b 2048 -ub 2048), is allowed to speculate, and arms the MTP head at n_max=1 when the model has one. We do not own a P40: this path is unmeasured, its status in the registry is INFERRED. The 11 GB 1080 Ti path is unchanged (checked on a 1080 Ti: -ub 768 and speculation declined, as before). Thanks to thisistimow for the P40 report.

PXA Control

  • KV is auto by default in the launcher and in PXA Control (no --ctk/--ctv unless you choose one).
  • GPU telemetry is included in the Report-a-problem bundle, with the same review-before-send rule as before.
  • A heat warning on the Rig and Launch pages when a card runs hot.

Checked on this build against v2026.10.1

Greedy 512-token output is byte-identical to v2026.10.1 on every Qwen row: 2x V100 (PXQN4), 2x P100 (PXQN4), 4x P100 (PXQN4), 1x V100 and 1x P100 (PXQN4 and the one-card file). Determinism 12/12 at np=1 and np=2 on the V100 pair and the P100 pair. 30k-token needle recalled on the V100 pair and on one P100. Decode is unchanged within noise on the rows this release does not touch: 1x V100 PXQN4 about 34 t/s, 2x P100 PXQN4 about 37.5 t/s, 1x P100 PXQN4 about 24 t/s.

Downloads

File For sha256
pxa-v2026.10.2-linux-x86_64-cuda12.8-sm60_61_70.tar.gz Ubuntu 24.04 and newer (glibc 2.38) 8ef2c50a3f9a44115df76ad5c3f60fdf99673793c186f3c484948e57aa07850f
pxa-v2026.10.2-linux-x86_64-cuda12.8-sm60_61_70-ubuntu22.04.tar.gz Ubuntu 22.04 and newer (glibc 2.35) 35d29b6e67664f890ec4daeabf63831b173d914cc104c9703630f6484a853b67
libggml-pxqn-v2026.10.2-linux-x86_64-cuda12.8-sm60_61_70.tar.gz PXQN library for source builds, Ubuntu 24.04+ 558e29add96fb2bd93845646fb6b0c046c11eb42892bea2de4294b8f80a89732
libggml-pxqn-v2026.10.2-linux-x86_64-cuda12.8-sm60_61_70-ubuntu22.04.tar.gz PXQN library for source builds, Ubuntu 22.04+ c220ab8c86f1a1e65eb04602f3186b8150e0a66c533b41c5b517c5f4641d5640

The PXQN library matches the tag v2026.10.2; see BUILD-FROM-SOURCE.md. The vLLM sidecar images are unchanged (pxa-vllm:sm60-v2026.10 / sm70-v2026.10).

Upgrade

Unpack over your old folder or into a new one. Same models.

Credits

mistrjirka, quenthalion, thisistimow (P40 report, first PXA-in-Odysseus confirmation).

Community

PXA v2026.10.1: every card, full speed

Choose a tag to compare

@poisonxa16 poisonxa16 released this 30 Sep 03:36

PXA Network

PXA v2026.10.1: every card, full speed

One card is 50% faster, Gemma speaks every language again, the 4x P100 crash is gone, and PXA Control now lets you report bugs and chase high scores.
Same models, no re-quant. Upgrade recommended.

Discord Ko-fi Docker

⚡ Qwen3.8-27B on v2026.10.1

Cards File Prefill Decode MTP prose MTP code
2× V100 PXQN4 924 56.4 77.1 107.6
2× P100 PXQN4 345 37.8 47.1 66.2
1× V100 One-card 1000 35.1
1× V100 PXQN4 1000 34.0
1× P100 One-card 252 24.2 32.3 36.2
1× P100 PXQN4 254 24.0
4× P100 PXQN4 440 30.7

📈 What the upgrade gets you

Setup v2026.10 v2026.10.1
1× V100 PXQN4 22.7 34.0 (+50%)
1× P100 PXQN4 19.9 24.0 (+21%)
2× P100 Ornith 35B ~60 ~67.6 (+12%)
4× P100, uneven VRAM crash runs
2× V100 / 2× P100 56.4 / 37.8 same

Tokens per second. Prefill is pp512 and decode is tg128 (llama-bench); MTP is the server's first answer to a fresh prompt, greedy, using the model's own MTP head (automatic on a tensor-split pair; one card: --spec-type mtp:n_max=1). Blank = not measured. Same file and flags on both builds. Output is byte-identical to v2026.10 on every row, and 12/12 determinism, 30k needle recall, MTP speed and English + Finnish output were rechecked on the final build.


🔧 What was wrong in v2026.10, and is now fixed

  1. Single-card and layer-split decode was about 30% slow. A fused delta-net kernel that only pays off with tensor split was on everywhere. It is now automatic: on for -sm tensor, off otherwise (PXA_DN_CONVFUSE still overrides).
  2. Gemma 4 26B-A4B broke on non-English text. The expert router was scaled twice. Fixed.
  3. 4x P100 PXQN tensor split crashed with uneven free VRAM. The split now lands on whole rotation blocks, so it boots and generates.
  4. Automatic n-gram speculation could collapse sampled chat (one user saw 80 → 19 t/s). Drafts are capped at 8 tokens, so sampled prose stays within a few percent of no speculation while edits still gain. PXA_SPEC_NGRAM_NMAX sets another depth.
  5. One V100 lost 5-8% to a GEMV fusion. It is now off on a single V100 for every tool, with identical output. The trade is about 2.5% slower prompt processing on a single V100, for the decode gain.

✨ New in PXA Control

  • 🐞 Report a problem. One click builds a report with your paths, host names, IPs and user names already removed, shows you every byte, and sends only when you press Send.
  • 🏆 Community high-score board. Benchmark your rig, then check the board for your model and card setup. Set a record and your name goes up in #benchmarks on our Discord. Nothing is sent without your click.
  • 💬 Discord and ♥ Support links right in the header.

🤗 Free showcase model: Qwen3.8-27B One-Card

poisonxa/Qwen3.8-27B-PXQN-OneCard: 12.6 GiB, 131k context on one 16 GB card, with its MTP head. On a single V100 it decodes 55.6 t/s prose / 61.2 t/s code with MTP (34 plain); on a single P100, 31.6 / 35.2 (23.6 plain). Needs PXA v2026.10.1.

🐳 Docker

docker pull ghcr.io/poisonxa16/pxa:v2026.10.1     # also :latest

Built from this release's 24.04 tarball (same binaries, byte for byte). The vLLM sidecars have no changes in this release: keep ghcr.io/poisonxa16/pxa-vllm:sm60-v2026.10 / :sm70-v2026.10.

⬇️ Downloads

File For sha256
pxa-v2026.10.1-linux-x86_64-cuda12.8-sm60_61_70.tar.gz Ubuntu 24.04 and newer (glibc 2.38) 08ffdbfe1bd12855f822a130b266ac862e9eab5bbc9ca26a1f3550c357d0d876
pxa-v2026.10.1-linux-x86_64-cuda12.8-sm60_61_70-ubuntu22.04.tar.gz Ubuntu 22.04 and newer (glibc 2.35) f13b8a2920a2cc3cd89058b21720d3de45333ed5c0b6f88dc4b01b76b257ffab

Both cover Pascal (sm_60, sm_61) and Volta (sm_70) with the CUDA 12.8 runtime bundled.

Building from source? The public source runs classic PXQ and K-quants as is. For PXQN files, grab the compiled PXQN library for your OS and drop libggml-pxqn.so next to your libggml.so (or set PXA_PXQN_LIB). It matches tag v2026.10.1 and gives the same output and speed as the tarball. See BUILD-FROM-SOURCE.md.

File For sha256
libggml-pxqn-v2026.10.1-linux-x86_64-cuda12.8-sm60_61_70.tar.gz source builds on Ubuntu 24.04+ b50d055c50cbe0bc7a0d5187290b75f2b0d58535314c26e79caaa865ee2e7492
libggml-pxqn-v2026.10.1-linux-x86_64-cuda12.8-sm60_61_70-ubuntu22.04.tar.gz source builds on Ubuntu 22.04+ 79ede0ad7dd3e6c14564187bba8b3f31f36fba450a40d8e287d629caa2983433

PXQN: closer to the original at a fraction of the size

PXQN: closer to the original at a fraction of the size. PXQN4 at 15.7 GB matches classic PXQ6 at 18.8 GB; PXQN5 matches Q6_K in 84% of the space. More on the project page.

Questions, rigs and benchmarks: discord.gg/EqazvV9tf · Keep the work going: ko-fi.com/shatteredrealms1

PXA v2026.10

Choose a tag to compare

@poisonxa16 poisonxa16 released this 29 Sep 12:48

Downloads

File For sha256
pxa-v2026.10-linux-x86_64-cuda12.8-sm60_61_70.tar.gz Ubuntu 24.04 and newer (glibc 2.38) 3dda43d0bc61c97e7b315e6a97ea943509a19a21d5ed264779de554d2cddd577
pxa-v2026.10-linux-x86_64-cuda12.8-sm60_61_70-ubuntu22.04.tar.gz Ubuntu 22.04 and newer (glibc 2.35) 74bc3c5951ebc2d6289fe812f39f32296cd98ffe7cf4ea40d58458817f9f0230

Both cover Pascal (sm_60, sm_61) and Volta (sm_70) with CUDA 12.8 runtime libraries bundled. The vLLM sidecar images ship separately.


PXA v2026.10 — release notes (draft)

On top of v2026.09.20.


Headline: PXQN, the next-generation quant format

PXA v2026.10 introduces PXQN, a rotated, Hessian-rounded evolution of the PXQ codec for
Pascal and Volta: a one-sided block-128 Hadamard rotation plus GPTQ/LDLQ-style Hessian-aware
rounding, built into a second-generation pxq-quantize. Old PXQ files keep loading; PXQN is a
new format id alongside them, not a replacement for them.

Speed next to size and quality — Qwen3.8-27B, plain decode (no drafting), 2026-09-28, final
binary.
Every speed figure below sits next to the file it was measured on: its size and its
KL divergence against the Q8_0 reference, scored on the assistant tokens of an in-distribution
chat set (what the model actually writes; user-turn tokens excluded).

2x P100, -sm tensor PXA, PXQN4 the fastest other Pascal build we tested, Q6_K
File size 15,720,262,976 B (4.25 bpw) 22,431,001,568 B (6.57 bpw)
KLD vs Q8_0, assistant tokens 0.007930 (same top token 97.06%) 0.002320 (same top token 98.28%)
Decode, tg128, t/s 37.85 32.54
Decode at 16k context, tg256, t/s 36.72 31.53
Prefill, pp512, t/s 344.95 273.67
Prefill, pp4096, t/s 337.58 272.10
Prefill, pp16384, t/s 316.68 267.15

llama-bench, both builds with the same flags (-ngl 99 -sm tensor -fa 1 -ctk q4_0 -ctv q4_0 -b 2048 -ub 512 -t 8 -r 3), same two cards, one bracket (ours / theirs / ours / theirs) after a 60 s
warm-up, judged on the closing pair, nothing else running on the box (prefill: -p 512,4096,16384 -n 0
in the same bracket). PXQN4 is a 30% smaller file than Q6_K, decodes 16% faster and reads a prompt 19
to 26% faster; Q6_K keeps the lower KLD — the trade is size and speed for a measured, small quality cost.

2x V100, -sm tensor, PXA file KLD vs Q8_0, assistant tokens decode tg128, t/s
PXQN4 15,720,262,976 B (4.25 bpw) 0.007930 56.00 (53.98 at 16k context)
PXQN5 18,785,381,696 B 0.002211 (lower than Q6_K) 48.80
Q6_K 22,431,001,568 B (6.57 bpw) 0.002320 48.32

PXQN4 row: final binary, quiet box, bracket close. PXQN5 and Q6_K rows and all KLD values: measured the same day on the
release-candidate binary one merge earlier (same decode kernels for these files at width 1).

MTP (the model's own draft head), first-pass numbers on the final binary. On a multi-card
-sm tensor run of a file that carries an MTP head, a bare llama-server now uses MTP by default:
draft depth 3 with a top-1 confidence floor (0.8 on P100, 0.9 on V100) that stops a draft chain
early when the head is unsure. Opt out with PXA_SPEC_AUTO_MTP=0 or --spec-type none.
Qwen3.8-27B PXQN4, greedy, 256 tokens:

t/s, first request of each class prose code
2x V100, plain 56.64 56.38
2x V100, MTP (bare server default) 77.14 107.71
2x P100, plain 38.46 38.52
2x P100, MTP (bare server default) 48.06 63.79
2x P100, the fastest other Pascal build we tested, Q6_K with its own MTP drafting 45.92 64.00

Fresh server per prompt class, prompt cache off, 60 s warm-up on an unrelated prompt, the first
request of each class quoted so nothing is replayed, three requests per cell, nothing else running.
Draft acceptance: 75.4% prose and 95.8% code on the P100 pair, 78.4% and 96.8% on the V100 pair.
On P100 our MTP is ahead of the other build's on prose and level on code (63.79 against 64.00).

Greedy output with MTP. MTP verifies several tokens in one wider step, and at that width the
recurrent (DeltaNet) layers, norms and attention round slightly differently from single-token
decode. Greedy text with MTP on is usually, but not always, byte-identical to MTP off: on this
binary both classes matched on the P100 pair, and code matched on the V100 pair, while prose
drifted there. This behaviour predates this release and is not a quality
loss; if you need byte-reproducible greedy output, leave MTP off.

PXQN in this release is compiled-only. The PXQN kernels ship as a separate library,
lib/libggml-pxqn.so, loaded next to libggml.so (RPATH $ORIGIN); the source tree carries only
the open forwarding layer. The release tarball includes it. A build made from the public source tree
runs every older format and PXQ, but it can load PXQN files only with the release build's
libggml-pxqn.so next to it.

vLLM images with PXQN. The vLLM sidecar images gain PXQN: PXQN4 and PXQN5 on V100
(pxa-vllm:sm70), PXQN4 on P100 (pxa-vllm:sm60), from pre-converted checkpoints
(Qwen3.8-27B PXQN4 and PXQN5). The images are leak-checked and are published
separately from the engine tarball.

Run MoE models that don't fit your cards. The expert cache keeps the hot experts on the GPUs
and serves the cold ones from host memory, planned automatically at load. It switches itself on
only for a MoE file that does not fit and has its <model>.expert-counts.csv next to it; every
file that fits runs exactly as before. Flash-Next on two P100s: 10.67 t/s decode and 188.6 t/s
prompt reading (3k tokens), with byte-identical output across runs.


What's new

Flash-Next ships with PXQN experts. Flash-Next (the qwen4exp MoE architecture) gets
LDLQ/Hessian-rounded PXQN experts alongside its dense PXQN Qwen3.8-27B sibling.
The shipped file is 98,660,237,024 bytes; its KL divergence against the reference on
assistant tokens is 0.0705 (same top token 91.2%), down from 0.0879 for the previous Flash-Next
file. On four P100s with the opt-in tensor-split recipe
(PXA_TSPLIT_QWEN4EXP_HC=1, see docs/LEVERS.md) it decodes 36.26 t/s (37.93 on prose) and reads
prompts at 510.3 t/s (3k tokens) and 474.4 t/s (16k tokens) on this release's final binary; the
launcher's default for this file stays the measured four-card layer recipe. Serve it from an
SSD, or set PXA_PLE_MMAP=0: memory-mapping its per-layer embedding table from a spinning disk
stalls the first long prompt.

Tensor split reaches PXQ/PXQN hybrid files. -sm tensor used to refuse every PXQ file with
a DeltaNet ssm_out tensor by name at load. A panel-aware row-range slicer
(PXA_TSPLIT_SSM_OUT_PANEL, on by default) fixes that for the dense Qwen3.8-27B PXQ/PXQN
tiers
on a matched two-card pair, and on four identical P100s. Bug #206 (wrong tokens on a 4-way split
inside a container with docker's 64 MiB /dev/shm) is fixed: a failed NCCL group now falls back
to the peer route, so 4x P100 takes the tensor split by default; PXA_TSPLIT_ALLOW_4WAY=0
restores the pair-only rule. This is not "every PXQ file, any card count" — see Known Issues
for what is explicitly excluded (Gemma 4 is a separate opt-in lever, not this default;
qwen35moe/Ornith fails its own KLD admission gate and stays on -sm layer; four V100s and
three-card sets are unmeasured and stay on -sm layer). Greedy tokens are not always identical
between -sm layer and -sm tensor under this slicer — see Known Issues.

Determinism. The speculative-decode cascade's n-gram stage now honors its configured depth
instead of building itself with an effective n_max of 4 regardless of what was asked for; this ships under the
existing default and changes nothing for a user who passes no flags. The 12/12-byte-identical
greedy gate holds on this release's final binary at np1 and at np2 with the second slot idle
(slot-pinned, same KV placement), together with the token-0 logit-spread check (6/6 identical)
and needle recall at about 3k, 11k and 21k tokens, for Qwen3.8-27B PXQN4 on the shipping defaults
of a P100 pair and a V100 pair.
Concurrent-decode nondeterminism (a second slot actively decoding a different prompt at the same
time) is not fixed this release — see Known Issues, it is not new to PXQN or to any file.
Server fixes. A handful of crash-on-a-bad-request bugs are closed: an unbounded
n_probs/top_logprobs value used to reserve that many entries per token on the decode thread
and throw bad_alloc outside any try block, killing the server (now capped at 1024); /slots/N save/restore/erase could hang forever on error or
success (server-slot-actions-hang-forever); an uninitialised stop field on a slot could throw
out of the decode loop during chat-message parsing (server-slot-stop-uninitialised-chat-parse);
and the streaming handler could hang forever when the very first result was already the final
one (server-stream-hangs-when-first-result-is-final). Parallel tool calls: a first fix for
the <tool_call> marker match (exact-string compare) shipped with a regression that dropped
parallel tool calls from streamed content; the correction (3a41a3bc1d, bug #222) is merged
into this release's integration branch. Prompt cache off by default on recurrent/hybrid
models
: CONFIRMED — the RAM prompt cache used to restore a similar-but-not-identical prior
prompt against a hybrid/recurrent model's state well enough to lose native tool-calling (bug
seat-ram-prompt-cache-degrades-hybrid); it is now off by default for models with recurrent
state unless --cache-ram is given explicitly (commit 5b4e83bbd3, confirmed an ancestor of
this release's integration commit dcb2b9dad7 — verified directly via git merge-base --is-ancestor, 2026-09-27).

The launcher picks its own settings, and so does the engine. A bare `llama...

Read more

PXA v2026.09.20

Choose a tag to compare

@poisonxa16 poisonxa16 released this 21 Sep 20:54

Known issue, fixed on main (2026-09-25): 4-card tensor split. On 4 identical GPUs the launcher's --sm auto picked -sm tensor for Qwen3.8-27B PXQ4, and the 4-way tensor split serves wrong tokens (perplexity looks normal, generation does not). Pairs are not affected. Fix: pass --sm layer on 4 cards, or take the updated tools/pxa-launch.py from main (commit d40e94c), which keeps the layer split beyond a pair. Tracked as bug #206.

Download: pxa-v2026.09.20-linux-x86_64-cuda12.8-sm60_61_70.tar.gz below (Linux x86_64, CUDA 12.8 runtime bundled, Pascal sm_60 / sm_61 and Volta sm_70).
sha256 a69a2dc7f427541692e2bc4b518998c26c9fe5772280ebe00c1aada0affa465e

tar xzf pxa-v2026.09.20-linux-x86_64-cuda12.8-sm60_61_70.tar.gz && cd pxa-v2026.09.20 && ./pxa-launch

Docker: docker pull ghcr.io/poisonxa16/pxa:v2026.09.20
Writing PXQ files: the quantizer is a separate download — https://github.com/poisonxa16/pxq-quantize/releases
Community: Discord https://discord.gg/EqazvV9tf · Ko-fi https://ko-fi.com/shatteredrealms1


PXA Network

PXA v2026.09.20

The first release that is not an RC: an inference engine for Tesla P100, V100 and GTX 1080 Ti, with its own quant format (PXQ) and its own kernels.

What's new

  • Faster by default on two matching cards. The launcher now picks a tensor split by itself. Qwen3.8-27B on 2× V100: 68.6 / 107.4 / 43.6 t/s (prose / repetitive text / ~15k-token prompt), up from 51.0 / 85.7 / 35.5. It falls back on its own when a model isn't proven on it, and --sm layer switches it off. Trade-off: prompt processing is slower on narrow PCIe links such as x4 risers.
  • Long context on V100. A new attention kernel: plain decode at 15k context 31.9 → 35.5 t/s, speculative decode 41.6 → 47.8 t/s.
  • Gemma 4 26B-A4B (the 128-expert model) works out of the box, on the stock file and on PXQ4 / PXQ3.
  • q8_0 KV cache: prompt processing on V100 is 16.7% faster.
  • The PXQ quantizer is a separate small download (CPU only): https://github.com/poisonxa16/pxq-quantize/releases — and it can now use an importance matrix for a closer file at the same size.
  • Boots and answers on a GTX 1080 Ti (checked with a small model).

Against stock llama.cpp — same cards, same window, each at its best settings (tokens/s)

this engine llama.cpp
Qwen3.8-27B, 2× V100, prose 80.2 58.9
… repetitive text 139.8 80.5
… ~15k-token prompt 47.8 46.1
4× P100, reading a 20.8k-token prompt 420 267
4× P100, decode 96.1 36.0

The long-context row is a narrow win; both engines' individual runs are printed in bench/LEADERBOARD.md. Rows where this engine runs another project's quant format are listed there too — it isn't tuned for those.

The numbers, as charts

PXA vs stock llama.cpp

Qwen3.8-27B on 2x V100, PXA vs mainline llama.cpp with its own MTP setup

The rows marked opt-in need a flag; plain decode is the out-of-the-box comparison. One more row is measured but left out on purpose: see the note under "The head-to-head, re-measured" below.

Tensor split on two identical cards

Earlier measurements

Prefill from the v2026.09.13 release candidate board, three engines on the same weights, and Gemma 4 measured on this package. Prefill never re-sends a prompt, so the correction below does not touch these numbers. The v2026.09.13 decode rows are left out on purpose: they were measured with repeated prompts, which the correction explains.

Prefill on 2x V100, three engines

Prefill on 4x P100, three engines, including the rows where mainline is faster

Gemma 4 26B-A4B on 2x V100, PXA vs stock llama.cpp


The full notes

On top of the v2026.09.13 release candidate.

The short version: out of the box, two matching cards now get the fused tensor split, and everything else new ships switched off, documented, with the number it measured and the command that turns it on. Gemma 4's 128-expert MoE is the other exception and ships on by default; its section has the gate table.

There is also a correction to publish about how I measured decode speed. It is at the top, not the
bottom.

The reference for every lever, what it does, its default and how to turn it off is
docs/LEVERS.md.


The correction — my speculative decode numbers were measured wrong

A speculative decoder keeps a table of text it has recently produced and uses it to guess what
comes next. My benchmark harness sent the same handful of prompts over and over inside one server
run, so from the second repetition onward the server was regenerating an answer it had already
written and the table could predict nearly every token of it. Acceptance climbed towards 100% and
the rate climbed with it — on one arm, 37 tokens/s on the first request and 120 by the sixth, for
the same prompt. Clearing the slot between repetitions does not clear that table.

Two honest things about it. It affected both sides of every comparison: the competitor ran
through the same harness in the same bracket, so the head-to-head shape was like-for-like even
where the absolute numbers were not. And the effect is real, just misnamed — an editor or a
tool-using client that regenerates text it has already seen genuinely does get that speed. What it
is not is the speed of a first answer, and a headline number should be a first answer.

Every decode figure in this note is re-measured:

  • a distinct prompt for every repetition, generated fresh, so no repetition is ever a
    regeneration;
  • three classes — free prose, a code-edit whose output re-quotes its own prompt, and a
    ~15,000-token long-context question — because a speculation change that helps one routinely
    hurts another;
  • the first-pass median is the headline;
  • six repetitions per cell where the number is a head-to-head claim, the competitor booted and
    driven in the same bracket on the same cards;
  • a plain control is really plain — an arm with no --spec-type is not a control on this
    engine, because the automatic rule arms an n-gram drafter, so controls pin it off;
  • every speculative cell names its acceptance rule, because a speculative rate without the
    rule beside it is not comparable to anything.

Prefill figures, the determinism and needle gates, and every lever A/B were unaffected: none of
them re-sends a prompt.


What changed for someone who changes nothing

Three fixes, plus one new default: on two matching cards the launcher now picks the fused tensor split (its own section below).

  1. The speculative verify path wrote back through the wrong tile. Checking a batch of draft
    tokens runs the model at width 2–8 instead of width 1. The decode kernel's write-back rule for
    that case clamped its tile backwards, so the widths speculation actually uses were the widths
    it handled worst. The fix is bit-identical by construction and it is on unconditionally.
  2. A prompt-cache bug that moved tokens it should not have. Reviewed, six findings fixed, and
    its test passes. It is unreachable in the default configuration anyway; it is fixed regardless.
  3. The drafter's cascade constants, including how often the n-gram table is fed. The table was
    allowed to lag up to 32 tokens behind the text; it is now fed every step, which is what the
    contract always said. Measured neutral on fresh text. PXA_SPEC_NGRAM_FEED_LAG=32 restores the
    old behaviour if you want to compare.

Two more are on only if their check passed on this binary — check the boot log for
PXA_AUTO: or PXA_PXQ_MMVQ_COLS lines to see which side of that this build landed on:
PXA_PXQ_MMVQ_COLS (a wider verify tile on Volta, engages only at verify width 3–8) and
PXA_MTP_ZERO_BUDGET_SKIP (a request that asks for no drafting stops paying for the drafter).

**Everyth...

Read more