Strix Llama 0.3.5: restored conversations keep all their memory; several conversations faster with MTP off
Strix Llama 0.3.5 fixes a bug that could blank part of a conversation's memory when it came back from the cache, and makes several conversations at once a little faster with MTP off.
A restored conversation keeps all of its memory. When a conversation comes back to a slot another conversation just used, from the memory cache or the disk cache, the server clears the cells the other conversation used and writes back the returning conversation's saved state. The clearing ran in the GPU's queue while the restore wrote from the CPU, and the last part of the clearing could land after the restore, blanking the returning conversation's values in the last attention layer for up to a few hundred positions. It was never another conversation's data: those positions came back empty. But the model lost some of that conversation's context in one layer, which could change its answers. It has been there since 0.1.17. The restore now waits for the clearing to finish, and a conversation's state as it left and as it came back is byte-identical. We found it while checking llama.cpp #29092, a report of recurrent state crossing between requests on this GPU. That one does not happen here: documents written in made-up vocabularies, answered one after another in the same slot, never had a word of another document in their answers.
Faster with several conversations, MTP off. Eight conversations of about 40K tokens each answering together in a 512K pool: 75.8 to 77.9 tokens a second with greedy decoding, and 73.2 to 77.7 with typical sampling settings, where the eight samples of each step are now taken in parallel instead of one after another. The sparse attention's block scoring also skips the blocks a conversation cannot see. One conversation alone runs as before, and the output is exactly the same as 0.3.4's.
Checked before release. Identical text and probabilities to 0.3.4 with MTP off, with MTP on, and with a q8_0 512K pool on eight slots; the new paths run beside the old ones in check modes over thousands of steps with no difference; a conversation's full state hashed as it leaves a slot and again when it comes back; and the usual checks: long answers with MTP, agents with regenerates and stopped answers, requests waiting for room in a full pool, images in several conversations, conversations read back from disk, the answer edges, and a 512K pool with conversations moving to disk and back.
Known limits.
- A conversation restored from the disk cache gets its system prompt's cached values as the first conversation with the same system prompt computed them, so at temperature 0 it can occasionally answer a late token differently than if it had never left the slot.
- A conversation with an image in it keeps the slower path, and the first image after the model loads takes longer (its kernels load on first use).
- Many conversations at once share 24 snapshots of the model's running state (about three each with eight), so going back far into an older part of a conversation can process more than it would alone.
- A request waiting for room in a full KV pool receives nothing until it starts, so a client that gives up quickly on a silent request can time out first.
- The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.