Strix Llama 0.3.2: answers start sooner, instant regenerate, less disk writing
Strix Llama 0.3.2 makes agents and multi-turn chats start answering sooner, regenerates an answer instantly, and writes less to disk.
Checkpoints where answers begin and end, not extra passes. This model keeps part of its memory as a running state that cannot be rewound, so the server saves that state at points a later request may need to go back to. Until now it saved them by cutting each prompt into extra pieces: a few tokens before its end, and at the start of its last message. Every cut is one more pass through the model, and when several agents send requests at once, each pass waits behind their work: the last four tokens of a prompt could wait two to three seconds on their own. The server now saves the state where an answer starts and where the previous answer ended - points it passes anyway - and cuts nothing. They cover everything the cuts covered: regenerating an answer, editing the last message, a client that drops the model's reasoning from the history, a tool call the client sends back reformatted, an answer stopped halfway. A main agent with five sub-agents on eight slots: the median wait for the first token went from 5.6 to 3.0 seconds, and the time spent on prompts from 139 to 87 seconds. The longest waits, when five sub-agents send their next step at the same moment, move less (9.6 to 8.3 seconds at the 90th percentile): they are bound by how fast the prompts themselves are processed.
Regenerate is instant. The point where an answer starts keeps the probabilities of its first token, so a regenerated answer starts from them at once without processing anything (0.3.1 processed four tokens).
Requests are served in the order they arrive. Prompts waiting for the model used to be taken in slot order, so a request that came in two seconds earlier could still wait for a later one.
Less disk writing for chats that drop reasoning. With Keep conversations on disk on, the server now writes only what a conversation is known to keep: its prompts, and an answer once the next request shows the answer is kept. A client that drops the reasoning from its history re-renders every answer, and such answers are no longer written. Six conversations with thinking on taking turns in two slots: 10.8 GB written before, 8.3 GB now, with no extra processing when they come back.
Checked before release. The full release checks again: identical text and probabilities to 0.3.1 with the new checkpoints switched off (with them on, the last few prompt tokens are computed together with the rest of the prompt, as they are for everything else, which moves the first answer token's probabilities slightly), perplexity identical over 40 chunks, long answers with MTP, agents on four slots with regenerates, edits and stopped answers, rewinds at 60K and 150K characters, images in several conversations, conversations read back from disk, the context's edges, and a 768K-cell KV pool with MTP.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
- The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.