Strix Llama 0.2.2
Chats get to their first token sooner.
No first-use stalls in a chat turn. hipBLAS loads each GEMM kernel from disk the first time it meets a shape. Two of the model's products still went through it:
- The sparse-attention indexer's BF16 projections, in any batch of 17 to 511 tokens. That is the new part of most chat turns, and it cost about 0.4 s the first time.
- The MTP draft head's Q6_K projection, in the draft's first prompt batch. That cost about 0.2 s.
Both now stay on this fork's own kernels.
A five-turn chat on a fresh server, two runs of each version, time to first token:
| 0.2.1 | 0.2.2 | |
|---|---|---|
| first message (3K tokens) | 4.3-4.9 s | 3.5-4.0 s |
| second turn (~570 tokens) | 1.5-1.6 s | 1.1-1.2 s |
Later turns vary with the answers.
Also in this release
- Output:
- A prompt that arrives as one batch of 512 tokens or more gives bit for bit 0.2.1's output.
- The smaller batches of a chat turn are summed in another order, so an answer can differ from 0.2.1's where two tokens were nearly tied.
STRIX_MMB_BF16_MIN_T=512 STRIX_MMQ_Q6K_ANY=0in the server's environment restores 0.2.1's routing.
- Output check:
- The 18.6K-token equivalence probe is bit for bit 0.2.1's.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity at 8K context is 2.6814 against 2.6811. The perplexity tool asks the output projection for 8192 rows at a time, which now run on MMQ; with
STRIX_MMQ_Q6K_ANY=0it gives 2.6811. The server asks for 1-16 rows, which MMQ already handled.
Install over 0.2.1; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 52-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/chat-ttft-20260925.json.
Requirements are unchanged:
- Ryzen AI Max+ 395 (gfx1151)
- Windows 11
- a 96 GB GPU carve
- the model files from
unsloth/Qwen3.8-Flash-Next-GGUF(UD-IQ4_XS) - the draft head asset from 0.1.2, if you want the last 5–9% of decode