You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Faster long-context prefill: FlashPrefill V2 reached 632.1 pp/s at 16K and 604.9 pp/s at 32K on Qwen3.8-27B Q4, R9700/HRX/LOOM. Default on for supported configurations; use --FlashPrefillV2=off for dense prefill.
MTP and DFlash2 support: at 16K, prompt throughput reached 582.6 pp/s with MTP3 and 615.2 pp/s with DFlash2. Both on/off pairs matched the 64-token output and aggregate acceptance. Drafting and verification stay dense. Includes approximately 5× faster block selection and Q4 SwiGLU bias-tail unrolling.
Previously observed DFlash2 peaks:67.6 tok/s decode and 96% acceptance. These are separate workload peaks; the new sparse tests establish prefill gains, not a decode speedup. Perplexity is not yet measured.