Strix Llama 0.3.4: faster answers with MTP off, flat memory with many conversations
Strix Llama 0.3.4 makes answers faster with MTP off, most of all with several conversations at once and long contexts, and stops the server's memory from growing with every turn.
Faster answers with MTP off. With MTP off, a 512K-token KV pool fits beside the larger models, which makes it the better setting for many conversations at once, so this release works on the time each token costs without MTP. Three things got cheaper. The model's sparse attention first scores every block of the context to pick the ones to read; that scoring was about twenty small passes over the GPU and is now one, reading each block once. Before every token, the server worked out which parts of the KV pool each conversation may see by walking the whole pool on the CPU while the GPU waited, which cost the most with several conversations; a text conversation is one unbroken run in the pool, so it is now described from where it starts. And some per-token host work is now done once and kept. Eight conversations of about 40K tokens each answering together in a 512K pool: 58.8 to 75.7 tokens a second in total. One conversation at 110K tokens: 21.6 to 24.6 tokens a second; the slowdown as a conversation grows is less than half what it was. With MTP on, the same work goes from every pass: a long answer at 79K tokens went from 33 to 35 tokens a second. The output is exactly the same as 0.3.3's, with MTP on or off.
Memory stays flat with many conversations. The server saves snapshots of the model's running state, about 113 MB each, at points a later request may need to go back to, up to eight per conversation. With eight conversations they added about 2 GB with every turn. All conversations now share a budget of 24 snapshots (about 2.7 GB); past it, the conversation holding the most gives up its least useful one, never the one at the end of its system prompt. Eight conversations taking three turns: 8.3 GB before, 4.8 GB now, with nothing processed again (a main agent with five sub-agents did the same work as before). One conversation on its own still keeps eight.
Checked before release. Identical text and probabilities to 0.3.3 with MTP off, with MTP on, and with a q8_0 512K pool on eight slots; both new paths were also run beside the old ones, comparing every result, across eight conversations, MTP, agents and a full pool (over 20,000 comparisons, no difference). Plus the usual release checks: perplexity over 40 chunks, long answers with MTP, agents with regenerates, edits and stopped answers, rewinds at 60K and 150K characters, images in several conversations, conversations read back from disk, requests waiting for room in a full pool, the context's edges, and a 768K-cell pool with MTP.
Known limits.
- A conversation with an image in it keeps the previous, slower path for both the scoring and the visibility work, and the first image after the model loads takes longer (its kernels load on first use).
- Many conversations at once share the 24 snapshots (about three each with eight), so going back far into an older part of a conversation, such as an edit several turns up, can process more than before.
- A request waiting for room in a full KV pool receives nothing until it starts, so a client that gives up quickly on a silent request can time out first.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.