Strix Llama 0.2.7
Several agents at once: agent sessions keep their conversations, eight stay loaded, and a long prompt no longer holds the others for many seconds at a time. Six agents on one system prompt, 0.2.6 with its four slots against 0.2.7 with its eight: tokens processed 75K → 45K, median time to the first token 10.1 → 3.4 s, the longest pause of a streaming agent 6.2 → 2.7 s.
Agent sessions keep their place. Agent tools run a session per sub-agent, all on one long system prompt, so any two sessions share most of their tokens. The server gave a new session whichever slot shared the most with it, often another agent's, and cut that agent's history, which then had to be processed again on its next step. A slot now takes a request for what it holds only when the request continues that conversation; anything else goes to a free slot and copies the shared system prompt on the GPU.
Eight conversations stay loaded. Concurrent conversations is now 8 by default, and can go up to 16. Six agents on four slots kept pushing each other out: nearly every step went back to the system prompt. With eight they stay put. Each conversation beyond the first costs about 0.4 GB of GPU memory, and a conversation in use keeps up to about 0.9 GB of snapshots in RAM. A configuration saved with an earlier version keeps its value: raise it there if you run agents.
One pass for all of them. When conversations were answering while new prompts came in, or several prompts arrived at once, the model made one pass over all its weights for each distinct prompt length: five sub-agents starting together took six passes. It is one pass now.
A long prompt no longer holds the others for long. While conversations are answering, a step now takes at most 2048 tokens of a new prompt. A 19K-token prompt used to hold a streaming conversation for 7 s at a time; now it pauses it for about 2 s at a time, and the long prompt itself takes about a fifth longer.
Update progress. The update prompt and the sidebar now show how far a download has got.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path.