Skip to content

Strix Llama 0.2.3

Choose a tag to compare

@nvwaonline nvwaonline released this 25 Sep 20:26
· 12 commits to main since this release

Several conversations at once are faster: three conversations decoding together get 17% more tokens, four 9% more.

A cheaper verify step for several conversations. With speculative decoding, each step verifies the drafted tokens of every conversation that is generating: three conversations drafting two tokens each verify nine tokens a step. At that size two of the model's products fell off their fast kernel:

  • The MoE router, from 9 tokens on. It ran on tiles of 128 rows, four workgroups for the whole GPU, about 250 µs a product.
  • The routed experts, from 5 tokens on. They ran on a tiled kernel that stages a tile per expert for the one or two tokens each expert gets.

Both now run the fast kernel over chunks of the batch. A nine-token verify step takes 103 ms instead of 120.

Drafts for four conversations. With the cheaper step, four conversations drafting two tokens each beat not drafting: 62.6 against 55.7 tok/s summed at ~4K tokens of context each, 56.1 against 50.6 at ~20K. The draft is now 3 tokens for one conversation, 2 for two to four, and none from five on. Six and eight conversations still do better without drafts.

ROCm 10.2. The runtime is built against TheRock ROCm 10.2.0a20260925. On the same source it gives the same perplexity as 10.1 and the same output bit for bit on the 18.6K-token equivalence probe. Prefill and single-conversation decode are unchanged; three or four conversations get 3-4% more.

0.2.2 against 0.2.3, alternating, three passes of two rounds each. Each conversation has ~4K tokens of context and generates 512 tokens with the server's default sampling. Means of six rounds:

0.2.2 0.2.3
one conversation 38.0 tok/s 37.5
three conversations 47.7 tok/s summed (1.26× one) 55.7 (1.47×)
four conversations 55.4 (1.46×) 60.4 (1.59×)

One conversation runs the same code as before; the 0.5 tok/s is one slower pass at the same 63% draft acceptance. With three and four conversations the acceptance depends on the sampled text, so single rounds spread by up to 5%.

Also in this release

  • The Strix Llama settings page has a one-line credit at the bottom.
  • Output:
    • One conversation's decode, and any batch of more than 32 tokens, give bit for bit 0.2.2's output.
    • A batch of 5-32 tokens is summed in another order: the verify step of several conversations, a short chat turn, the last few tokens of a prompt. An answer can differ from 0.2.2's where two tokens were nearly tied.
    • STRIX_F32_VEC_CHUNK_MAX=0 STRIX_MOE_VEC_CHUNK=0 in the server's environment restores 0.2.2's routing.
  • Output check:
    • The 18.6K-token equivalence probe is bit for bit 0.2.2's.
    • The long-output test reaches all 30 sections twice.
    • A 40-turn agent workload on four slots logs no warning.
    • Perplexity at 8K context is 2.6814, as in 0.2.2.

Install over 0.2.2; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 52-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/multi-stream-20260926.json.

Requirements are unchanged:

  • Ryzen AI Max+ 395 (gfx1151)
  • Windows 11
  • a 96 GB GPU carve
  • the model files from unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ4_XS)
  • the draft head asset from 0.1.2, if you want the last 5–9% of decode

Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.