Skip to content

Strix Llama 0.2.4

Choose a tag to compare

@nvwaonline nvwaonline released this 26 Sep 05:06
· 12 commits to main since this release

A small release for several conversations at once: a verify step of three or four conversations is 3-5% shorter, which gives 2-4.5% more tokens. It also corrects the numbers 0.2.3 published.

Three more small products off few-workgroup kernels. With speculative decoding, three conversations drafting two tokens each verify nine tokens a step. After 0.2.3, three products of that step still ran on kernels that gave the whole GPU only a few workgroups:

  • The input of the gated delta net layers' convolution, which appends each conversation's new tokens to its conv state. Below 32 tokens it took a generic copy kernel with one block per channel and conversation, about 30,000 small blocks at three conversations. It now takes the tiled transpose from 2 tokens.
  • The sparse-attention indexer's BF16 key projection, from 9 tokens on. It ran on one 128-row tile: one workgroup.
  • Quantized weights with at most 1024 output rows, from 9 tokens on: the hyper-connection down projection (96 times a step), the attention keys and values, the shared expert. Three to five tiles each.

The last two now run the vector kernel over chunks of 8 tokens. By the server's own timing over ~1,900 steps per build, a nine-token verify step takes 103.0 ms instead of 106.4, a twelve-token one 123.5 instead of 129.4.

0.2.3 against 0.2.4, alternating, three passes of three rounds. Each conversation has ~4K tokens of context and generates 512 tokens with the server's default sampling, with a new seed every round (the same seeds for both builds). How many drafts are accepted depends on the text a round happens to sample (54-86% here), and that moves tok/s more than this change does, so the table compares the builds at the same acceptance:

0.2.3 0.2.4
three conversations, at 68% acceptance 53.8 tok/s summed 55.0 (+2.2%)
four conversations, at 67% 58.1 60.7 (+4.5%)
one conversation 37.5 37.6 (the same output bit for bit)

A correction to 0.2.3's numbers. The 0.2.3 notes said three conversations got 17% more tokens than with 0.2.2, and four 9% more. That comparison fixed each conversation's sampling seed. A server that computes every step the same way then samples the same text in every round, and the acceptance is that text's: 0.2.2 sampled nearly the same text every round (66-67%), 0.2.3 did not (65-72%), so its average was partly the luck of the texts it drew. At the same acceptance, 0.2.3 got 13% more than 0.2.2 at three conversations, and 8% more at four.

Also in this release

  • With fixed seeds, three conversations now repeat their text from run to run: the same texts in four runs out of four, where 0.2.3 gave three different outcomes in four. Up to 0.2.3 these products changed kernel at 9 tokens, so a token's result could depend on how many tokens its step happened to carry. Four conversations can still vary.
  • Output:
    • The decode of one or two conversations, and any batch of more than 32 tokens, give bit for bit 0.2.3's output.
    • A batch of 9-32 tokens is summed in another order: the verify step of three or four conversations, a short chat turn, the last few tokens of a prompt. An answer can differ from 0.2.3's where two tokens were nearly tied.
    • STRIX_BF16_VEC_CHUNK_MAX=0 STRIX_Q_VEC_CHUNK_MAX=0 in the server's environment restores 0.2.3's routing.
  • Output check:
    • The 18.6K-token equivalence probe is bit for bit 0.2.3's.
    • The long-output test reaches all 30 sections twice.
    • A 40-turn agent workload on four slots logs no warning.
    • Perplexity at 8K context is 2.6814, as in 0.2.3. With 16-token batches, where every batch takes the new paths, it is 4.0732 against 4.0870 with 0.2.3's routing.

Install over 0.2.3; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 53-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/verify-small-products-20260926.json.

Requirements are unchanged:

  • Ryzen AI Max+ 395 (gfx1151)
  • Windows 11
  • a 96 GB GPU carve
  • the model files from unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ4_XS)
  • the draft head asset from 0.1.2, if you want the last 5–9% of decode

Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.