You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Faster attention: share K/V reads across query heads and token rows during speculative verification.
Automatic Loom optimization: enable loop-invariant code motion, with cooperative staging for FP16, BF16, FP8 and BF8.
Real-world peak performance:67.6 output tokens/s, 390.2 prompt tokens/s, and 96% draft acceptance with Qwen3.8-27B Q4 + Q8 DFlash2 on AMD R9700, macOS HRX/Loom. Peaks from individual requests before the reported tool-call loop.