Strix Llama 0.2.1
Several conversations at once: with speculative decoding on, three conversations decoding together are about a third faster.
Drafts of one length. With several conversations generating, the MTP drafter stopped each conversation's draft at its own first unconfident token, so one step could draft 2, 2 and 1 tokens. This model's hybrid memory processes a batch in pieces with the same number of tokens for every conversation. An uneven verify therefore ran as two or three passes of the whole model, each in a shape the GPU graph cache had not seen: 200 to 440 ms, where an even one takes about 120. Three conversations at once got 34 tok/s summed with drafting on, fewer than the 46 they got with it off.
Now the drafts of one step all have the same length. A conversation that turns unconfident keeps drafting while another is still confident, and all drafts are cut to the longest confident run.
| before | now | |
|---|---|---|
| three conversations, ~4K tokens each | 34.3 tok/s summed | 45.7 |
| three conversations, ~20K tokens each | 39.6 | 45.7 |
| two conversations | swung between 26 and 41 | 39-42 |
One conversation is unchanged, and so are four or more, which do not draft.
No 270 ms stall when several conversations start decoding together. For several conversations, sparse attention computes which conversation each block of the cache belongs to. That product went to hipBLAS, which loads a kernel from disk the first time it meets a shape: 272 ms in the first step three conversations decoded together, and again as their contexts grew. It now has its own small kernel (0.02 ms, same result).
Also in this release
- Corrected: 0.2.0's notes gave 0.1.17's perplexity as 2.6880, but that was measured on a development build. On the 0.1.17 release it is 2.6841, against 0.2.0's 2.6811. On the older test (English and code, 4K context) it is 2.4741 against 2.4730. The accuracy is the same.
- Where the time of a step with several conversations goes, and why gufo's multi-user table is not comparable: every user there is sent the same prompt, so all of them route to the same experts and the weights are read once. See
docs/results.md. - Output check:
- One conversation's output is bit for bit 0.2.0's.
- With q8_0, a conversation processed over another's freed cells hashes the same as alone after every prompt batch.
- The same 85K-token prompt sent twice gives the same 400 tokens as 0.2.0.
- Two long conversations (173K and 155K tokens) read back from disk answer exactly as when both stay resident.
- The image stress test passes with the disk tier on.
- A 137K-token conversation with short prompts in between logs no assert.
- A server that failed an allocation at a checkpoint answers the retry exactly as one that never failed.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity is unchanged.
Install over 0.2.0; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 51-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/multi-stream-20260925.json.
Requirements are unchanged: Ryzen AI Max+ 395 (gfx1151), Windows 11, a 96 GB GPU carve, the model files from unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ4_XS), and the draft head asset from 0.1.2 if you want the last 5–9% of decode.