Skip to content

LSE v0.4.21 — joint decode attention and automatic LICM

Choose a tag to compare

@Geramy Geramy released this 30 Sep 05:53
  1. Faster attention: share K/V reads across query heads and token rows during speculative verification.
  2. Automatic Loom optimization: enable loop-invariant code motion, with cooperative staging for FP16, BF16, FP8 and BF8.
  3. Real-world peak performance: 67.6 output tokens/s, 390.2 prompt tokens/s, and 96% draft acceptance with Qwen3.8-27B Q4 + Q8 DFlash2 on AMD R9700, macOS HRX/Loom. Peaks from individual requests before the reported tool-call loop.