Skip to content

Strix Llama 0.3.3: requests wait instead of failing when the KV pool is full; total throughput

Choose a tag to compare

@nvwaonline nvwaonline released this 28 Sep 08:35
· 8 commits to main since this release

Strix Llama 0.3.3 lets requests wait for room when the conversations together need more KV cache than the pool holds, instead of failing them, and shows the server's total throughput on the Logs page.

Requests wait instead of failing when the KV pool is full. With several conversations at once - agents especially - the KV pool can hold less than they need together. Until now, when a step no longer fit, every request running at that moment failed with "Context size has been exceeded", even the ones already answering, and their cached conversations were dropped. Now a request that does not fit waits in the queue until there is room: idle conversations make room first (they go to disk when Keep conversations on disk is on), and running ones finish. Running answers get room before new prompts. Only if every running conversation is stuck waiting for room does the one that arrived last give way, with an error a client can retry. Eight new 12K-token conversations sent at once to a 64K pool: 0.3.2 failed all eight, 0.3.3 answered all eight (three of them after waiting). The Logs page shows how many requests are waiting.

Throughput across all conversations. Next to Conversations on the Logs page, a line now shows how many answer tokens and prompt tokens a second the server is producing, summed over every conversation and averaged over the last 10 seconds. With several agents or chats running at once, a single conversation's speed says little; this is what the machine is doing in total. It comes from the same conversation counters the page already reads, so it costs the server nothing.

Known limits.

  • A request waiting for room receives nothing until it starts, so a client that gives up quickly on a silent request can time out first. A KV pool at least as large as the conversations you run at once avoids the wait.
  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.