Strix Llama 0.3.1: a used slot's block keys, cheaper MTP drafts at depth
Strix Llama 0.3.1 fixes a quiet quality bug in long conversations and makes MTP decoding cheaper deep into a long context. Updating is recommended: the bug is in every earlier release, and the default eight conversation slots make it easy to hit.
Where it stands, at the default settings. A 95.6K-token prompt prefills at 1228 tokens a second; answers decode at 36.7 tokens a second after 86K tokens of context and 44.4 after a short question (MTP on); three and four conversations at once decode 55.4 and 62.5 tokens a second together (each step takes the same time as in 0.3.0; the sum moves with how many draft tokens get accepted).
A new conversation in a used slot reads its own beginning again. The sparse attention keeps a summary key for every block of 4 tokens. A new conversation that started in a slot another conversation had used, with a first message under about 2K tokens, kept the old conversation's summary keys for its first blocks. Once the conversation grew past 2K tokens, the model's choice of which earlier parts to attend to was scored against the wrong keys for its system prompt and first message. The answers were plausible, just not the model's: the same conversation gave different answers on a freshly started server and after another conversation in the slot. Now both give the same answer, token for token.
MTP drafts are cheaper deep into a conversation. The small draft model that proposes the next tokens read its whole context at every step. It now reuses the part of the context the model picked as relevant a few tokens earlier, as SGLang and TensorRT-LLM do for DeepSeek's sparse attention, and refreshes that choice every 32 tokens. A step of decoding (a model pass plus its drafts): 3.6% faster at 86K tokens of context, 6.5% at 212K, with the same share of drafts accepted. Below 32K tokens nothing changes. The draft also no longer computes attention for tokens it only stores, which trims a little from every step.
Checked before release. The full release checks of 0.3.0, again: text and probabilities identical to 0.3.0 with MTP off, perplexity identical chunk by chunk over 40 chunks, long answers with MTP, agents on four slots, rewinds with MTP at 60K and 150K characters, images in several conversations, conversations read back from disk, the context's edges, and a 768K-cell KV pool with MTP.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
- The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.