Skip to content

Strix Llama 0.3.0

Choose a tag to compare

@nvwaonline nvwaonline released this 27 Sep 19:43
· 2 commits to main since this release

Strix Llama 0.3.0 is the stable release of the 0.2 series: it fixes the last known way a setting could keep the model from loading, trims a little host time when several long conversations answer at once, and went through the full release checks plus a long mixed workload before release.

Where it stands, at the default settings. A 95.6K-token prompt prefills at 1217 tokens a second; answers decode at 35.8 tokens a second after 86K tokens of context and 44.9 after a short question (MTP on); three and four conversations at once decode 58.8 and 62.6 tokens a second together.

Large KV pools with MTP load. With MTP on, a KV pool above about 512K cells failed to load: the draft reserved a working buffer over the whole pool (3.9 GB at 768K). The draft now works in smaller steps when the pool is larger, so pools up to the 1M-cell maximum load. A q8_0 pool of 768K cells fits in GPU memory at four conversations; larger pools spill into shared GPU memory, which works but is slower.

Several long conversations: less host time a step. Building the attention mask for a step that serves several conversations read every cell's bookkeeping once per conversation; it now reads it once. The result is byte for byte the same.

Checked for stability. Before release it ran 150 minutes of mixed work at the default settings - agents, images in several conversations, the prompt cache, long prompts, several conversations at once - 145 workloads without a failure, its memory flat after the first half hour. Prompts longer than the context are refused with a clear message, an answer that reaches the end of the context stops cleanly, and the server keeps answering afterwards.

Known limits.

  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.