Skip to content

Strix Llama 0.2.9

Choose a tag to compare

@nvwaonline nvwaonline released this 27 Sep 14:19
· 4 commits to main since this release

A small release: sparse attention now covers every size of prompt piece, and prefill is a little faster.

Sparse attention for pieces of 33-127 tokens. In a long conversation, a piece of a prompt between 33 and 127 tokens - typical agent tool results and short follow-up messages - used dense attention over the whole conversation instead of the model's sparse attention. Such pieces now use the sparse selection like every other size.

A guard against the 0.2.7 kind of bug. If sparse attention ever reaches a kernel that would ignore its selection, the server now stops with an error instead of computing a wrong answer.

Prefill a little faster. The recurrent layers' kernel no longer spills to scratch memory, and the per-head norms run a faster kernel; the results are bit for bit the same. A 95.6K-token prompt: 1226-1231 → 1242-1245 tokens a second.

Known limits.

  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path.