Strix Llama 0.2.8
A fix for 0.2.7: with several long conversations loaded, a prompt that arrived while others were answering could take minutes and be answered from a wrong computation. Prompts next to long answering conversations are also faster now.
What went wrong. 0.2.7 put a new prompt and the conversations that were answering into one pass. When the conversations loaded together held more than about 262,000 tokens - easy with a larger KV pool, with the default pool only when it is nearly full - the sparse attention's fast kernel could not take such a pass, and the fallback computed the prompt's attention over every token in memory, other conversations included. A few thousand tokens then took over a minute, the other conversations stood still meanwhile, and the conversation's answer came from a wrong computation (in our test it read differently from the correct one, though still fluent). The fast kernel now takes a pass of any size.
Stored conversations are cleared once. 0.2.7 may have kept such a conversation in Keep conversations on disk, so 0.2.8 empties that store on its first start. A conversation you come back to is processed again once.
A long prompt runs in a pass of its own when that is faster. Sharing a pass costs a prompt the other conversations' contexts, so next to long conversations a prompt now runs separately; short requests - answers in progress, agent tool results next to small conversations - still share one pass. Four slots with three long conversations answering: a 49K-token prompt 793 → 1035 tokens a second, a 74K-token prompt 695 → 1006, a 2.4K-token new chat 4.8 → 3.3 s; a conversation read back from disk with a 3.9K-token tail 76.6 → 4.3 s, the others pausing for 3 s instead of 41. GPU memory is unchanged.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path.