Skip to content

Strix Llama 0.2.6

Choose a tag to compare

@nvwaonline nvwaonline released this 27 Sep 01:57
· 8 commits to main since this release

Several conversations and agents without reprocessing: conversations stay loaded, a shared system prompt is computed once, and decoding several conversations at once is a little faster.

Four conversations stay loaded. Concurrent conversations is now 4 by default (Model › Configuration › Concurrency). Each conversation keeps its place in the model, so going back to one continues where it was instead of processing it again, and up to four answer at the same time. They share one context pool, so a single conversation can still use all of it. Each conversation beyond the first costs about 0.4 GB of GPU memory. A configuration saved with an earlier version keeps its value: set it to 4 there if it says 1.

Image input with several conversations. Image input used to need Concurrent conversations at 1: an image in a second loaded conversation stopped the server. That is fixed, and both work together.

Agents: a shared system prompt is computed once. Agent tools start many sessions with the same long system prompt and tool list. A new session now takes the part it shares from a session that already computed it, copied on the GPU in tens of milliseconds, instead of computing it again. Sessions that start at the same moment wait for the first one to get past that part, then copy it. An agent workload with a 10K-token system prompt, three sessions started together and three sub-agents: 29K tokens processed instead of 73K, prompt time 172 → 60 s, the three first answers after 15–17 s instead of 36–41 s.

Edits and regenerations reuse what they can. The model can resume a conversation only at the points where it kept a snapshot of its state. Those now sit where a later message is likely to branch: where the system prompt ends, at each of your messages, and at the end of each prompt. Editing an earlier message resumes from that message instead of from the start (the second message of an 18K-token conversation: 4.3K tokens processed before, 1.1K now). Regenerating and editing the last message were already cheap and still are.

Faster with several conversations at once. When several conversations decode together, each step checked the drafted tokens of all of them and saved the model's state after every checked token, 3 MB per layer, in case a draft was rejected. It now keeps only what it needs to recompute a rejected tail. Each step at three and four conversations is 3–4% shorter. Output is identical to 0.2.5.

Known limits.

  • More conversations than Concurrent conversations, with Keep conversations on disk off (the default): a conversation that had to make room is processed again when it comes back. Raise Concurrent conversations, or turn on Keep conversations on disk.
  • While one conversation starts or resumes, the others pause briefly: about 0.1–0.3 s, over a second when it is read back from disk. Removing that pause is the next release's work.