Strix Llama 0.2.4
A small release for several conversations at once: a verify step of three or four conversations is 3-5% shorter, which gives 2-4.5% more tokens. It also corrects the numbers 0.2.3 published.
Three more small products off few-workgroup kernels. With speculative decoding, three conversations drafting two tokens each verify nine tokens a step. After 0.2.3, three products of that step still ran on kernels that gave the whole GPU only a few workgroups:
- The input of the gated delta net layers' convolution, which appends each conversation's new tokens to its conv state. Below 32 tokens it took a generic copy kernel with one block per channel and conversation, about 30,000 small blocks at three conversations. It now takes the tiled transpose from 2 tokens.
- The sparse-attention indexer's BF16 key projection, from 9 tokens on. It ran on one 128-row tile: one workgroup.
- Quantized weights with at most 1024 output rows, from 9 tokens on: the hyper-connection down projection (96 times a step), the attention keys and values, the shared expert. Three to five tiles each.
The last two now run the vector kernel over chunks of 8 tokens. By the server's own timing over ~1,900 steps per build, a nine-token verify step takes 103.0 ms instead of 106.4, a twelve-token one 123.5 instead of 129.4.
0.2.3 against 0.2.4, alternating, three passes of three rounds. Each conversation has ~4K tokens of context and generates 512 tokens with the server's default sampling, with a new seed every round (the same seeds for both builds). How many drafts are accepted depends on the text a round happens to sample (54-86% here), and that moves tok/s more than this change does, so the table compares the builds at the same acceptance:
| 0.2.3 | 0.2.4 | |
|---|---|---|
| three conversations, at 68% acceptance | 53.8 tok/s summed | 55.0 (+2.2%) |
| four conversations, at 67% | 58.1 | 60.7 (+4.5%) |
| one conversation | 37.5 | 37.6 (the same output bit for bit) |
A correction to 0.2.3's numbers. The 0.2.3 notes said three conversations got 17% more tokens than with 0.2.2, and four 9% more. That comparison fixed each conversation's sampling seed. A server that computes every step the same way then samples the same text in every round, and the acceptance is that text's: 0.2.2 sampled nearly the same text every round (66-67%), 0.2.3 did not (65-72%), so its average was partly the luck of the texts it drew. At the same acceptance, 0.2.3 got 13% more than 0.2.2 at three conversations, and 8% more at four.
Also in this release
- With fixed seeds, three conversations now repeat their text from run to run: the same texts in four runs out of four, where 0.2.3 gave three different outcomes in four. Up to 0.2.3 these products changed kernel at 9 tokens, so a token's result could depend on how many tokens its step happened to carry. Four conversations can still vary.
- Output:
- The decode of one or two conversations, and any batch of more than 32 tokens, give bit for bit 0.2.3's output.
- A batch of 9-32 tokens is summed in another order: the verify step of three or four conversations, a short chat turn, the last few tokens of a prompt. An answer can differ from 0.2.3's where two tokens were nearly tied.
STRIX_BF16_VEC_CHUNK_MAX=0 STRIX_Q_VEC_CHUNK_MAX=0in the server's environment restores 0.2.3's routing.
- Output check:
- The 18.6K-token equivalence probe is bit for bit 0.2.3's.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity at 8K context is 2.6814, as in 0.2.3. With 16-token batches, where every batch takes the new paths, it is 4.0732 against 4.0870 with 0.2.3's routing.
Install over 0.2.3; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 53-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/verify-small-products-20260926.json.
Requirements are unchanged:
- Ryzen AI Max+ 395 (gfx1151)
- Windows 11
- a 96 GB GPU carve
- the model files from
unsloth/Qwen3.8-Flash-Next-GGUF(UD-IQ4_XS) - the draft head asset from 0.1.2, if you want the last 5–9% of decode
Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.