Repository navigation
Qwen3.8-Flash-Next on 6x3090 and 6x4090 without NVLink: prefill flat to 250k and faster, decode 2-10x vs upstream at depth, binaries and a reproducible grid #30071
lukolszewski
started this conversation in
Show and tell
Replies: 1 comment 1 reply
|
Can you fit latest MoE model GLM5.3 Flash in 2-3 or ideally 4 bit quant into your setup? Your speeds are mind blowing. When I ask AI if more RTX3090 cards will speedup my GLM5.3 Flash it tells that context size for sure will increase but most probably speed will be lower not higher. |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment

Uh oh!
There was an error while loading. Please reload this page.
I have been running Qwen3.8-Flash-Next (qwen4exp) on six RTX 3090s behind consumer PCIe (4 of them hanging off USB4 eGPU docks), five simultaneus sessions of up to 262k each, since September and fixing what made it slow. I've been very pleasantly surprised with the great speed increase with long contexts and multiple session use.
The patches sit on top of upstream df03399 and are published with binaries and container images here:
https://github.com/lukolszewski/llama.cpp-multigpu
It is still llama.cpp: same GGUF files, same llama-server API and flags, one build, every patch selected at runtime by an env var or a server flag. CUDA only.
Upstream's rate falls with context and collapses when several sessions decode at once; the patched build stays roughly flat. Same model (unsloth UD-Q4_K_XL), same command line on both sides, 5 slots x 262k, q8_0 KV, layer split over six GPUs. Upstream is df03399 built with the same recipe. Tokens/s, 5-session rows per session (more info in the README):
At 5k with one session there is nothing to gain (1.0x on the 4090s). The 10.8x is a 2.3 t/s baseline, so the shape of the curves is the claim, not the ratio at the far end. Full 20-row grids, hardware records, raw JSON and server logs are in the README and benches/.
Binaries support Tesla V100 and RTX2000 (as well as all the modern archs including RTX5090) as I'm plannign to test it on v100 (sadly I couldn't rent a working one today) and modded RTX2080 ti 22GB - If I can get my hands on a mashine with 6-8 of those. It is fun to be able to get very usable results from relatively budget over half decade old hardware, but the same code still provides benefits on newer GPUs.
To run it, take the release tarball (cuda-12.9: V100 to RTX 5090 on any R525+ driver; cuda-13.4: Ampere+) or the server image, and use the command line from README section 5. With one GPU no switches are needed(although I'm targetting multigpu). To reproduce the grid on your own box the bench image runs both sides, downloads the model and the upstream baseline itself, and writes the JSON the tables are generated from; one docker command, also in the README. If you run it on a 5090 or PRO 6000 box I would like the JSON and hardware.md back.
Two separate problems are fixed. The first is the decode slowdown with context depth from #28734: block-granular indexer top-k with a persistent pooled-key cache, and a compact-gather attention path so q8_0 KV keeps the sparse speedup. That part is built in, has no switch, and is what helps single-GPU users. The second is layer split over PCIe without NVLink: several sequences decoding without the GPUs waiting on each other, pipeline-parallel prefill, prefilling one slot while others decode, async prompt-cache transfers, and n-gram lookup speculation that works with all of that. Each change is one commit, listed with its switch in docs/multigpu/patches.md, and the series is attached to the release as format-patch output.
Caveats: one model family, CUDA only, written for slow PCIe, so NVLink or single-GPU systems may not benefit and some patches could regress them if turned on. Mixed prefill plus decode is better than upstream but still the weak spot. Only sm_86 runs every day; the 4090 rows are from one rental. The code was written with an LLM and validated by measurement and output checks, not by review, so I will not open PRs from it myself; anyone is welcome to pick up any piece. The repo exists to be archived once upstream catches up. Upstream has meanwhile reworked the indexer; my 2026-10-03 port onto master is in #28734 and was slower on this hardware.
I tried the MTP and it didn't help, slowed me down in most cases, so no MTP support in this branch, but n-gram speculation does 2.5x specifically code refactor for single user. I added an option to automatically disable it beyond 2 active decodes so we get the benefit without the downside, but our benchmarking did not not enable it to measure throughput without speculation.
Thanks to @svgop, @wussh, @goodbadwolf and @albertnsoliz for reproducing and analysing the depth decay on a 5090, an A100 slice and a PRO 6000 in #28734.
All reactions