veloGB10 v0.4.0
veloGB10 v0.4.0 — Qwen3.8 27B NVFP4 + DFlash 2, TP=4
Qwen3.8 27B NVFP4 support with native DFlash 2 speculative decoding, full 256K context, and TP=4 serving.
- New Qwen3.8 27B NVFP4-FULL model support, launched with
--spec-source dflash2-auto --draft-dir <dflash2 dir>. - TP=4 serving (three nodes + head), plus TP=2 and single-node.
- New DSV4 / DFlash2 / DSpark / MXFP4 kernel set (all
kernels/*.cu+ allsrc/ptx/*.ptx). - README: Update section with Qwen3.8 27B performance table and live throughput traces.
- New docs: QWEN_27B_SETUP.md and MANAGING_CACHE.md.
Performance (Qwen3.8 27B NVFP4 + DFlash 2, greedy, full 256K ctx)
| Mode | Average | Bottoms | Peaks |
|---|---|---|---|
| Single node | > 30 tok/s | ~18 tok/s | 45-50 tok/s |
| TP=2 | 56 tok/s | ~30 tok/s | 75 tok/s |
| TP=4 | 75 tok/s | ~45 tok/s | 125 tok/s |
Peak rates are typically on code content generation. See the README for the full performance traces.
Downloads
- velogb10-v0.4.0-gb10-sm121.tar.gz - the engine binary + all PTX kernel artifacts + TP launch scripts
- SHA256SUMS.txt - checksums
- PROVENANCE.txt - build provenance
Built from the public repo tree (self-contained), binary sha256 ef0007a533ee5f476aed3ae3930873b991e8190d323ae3fe5603f102cef6f90e.