Strix Llama 0.2.9
A small release: sparse attention now covers every size of prompt piece, and prefill is a little faster.
Sparse attention for pieces of 33-127 tokens. In a long conversation, a piece of a prompt between 33 and 127 tokens - typical agent tool results and short follow-up messages - used dense attention over the whole conversation instead of the model's sparse attention. Such pieces now use the sparse selection like every other size.
A guard against the 0.2.7 kind of bug. If sparse attention ever reaches a kernel that would ignore its selection, the server now stops with an error instead of computing a wrong answer.
Prefill a little faster. The recurrent layers' kernel no longer spills to scratch memory, and the per-head norms run a faster kernel; the results are bit for bit the same. A 95.6K-token prompt: 1226-1231 → 1242-1245 tokens a second.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path.