Skip to content

LSE v0.4.23

Choose a tag to compare

@Geramy Geramy released this 30 Sep 19:52
  1. Faster long-context prefill: FlashPrefill V2 reached 632.1 pp/s at 16K and 604.9 pp/s at 32K on Qwen3.8-27B Q4, R9700/HRX/LOOM. Default on for supported configurations; use --FlashPrefillV2=off for dense prefill.
  2. MTP and DFlash2 support: at 16K, prompt throughput reached 582.6 pp/s with MTP3 and 615.2 pp/s with DFlash2. Both on/off pairs matched the 64-token output and aggregate acceptance. Drafting and verification stay dense. Includes approximately 5× faster block selection and Q4 SwiGLU bias-tail unrolling.
  3. Previously observed DFlash2 peaks: 67.6 tok/s decode and 96% acceptance. These are separate workload peaks; the new sparse tests establish prefill gains, not a decode speedup. Perplexity is not yet measured.