Skip to content

Strix Llama 0.2.2

Choose a tag to compare

@nvwaonline nvwaonline released this 25 Sep 15:17
· 19 commits to main since this release

Chats get to their first token sooner.

No first-use stalls in a chat turn. hipBLAS loads each GEMM kernel from disk the first time it meets a shape. Two of the model's products still went through it:

  • The sparse-attention indexer's BF16 projections, in any batch of 17 to 511 tokens. That is the new part of most chat turns, and it cost about 0.4 s the first time.
  • The MTP draft head's Q6_K projection, in the draft's first prompt batch. That cost about 0.2 s.

Both now stay on this fork's own kernels.

A five-turn chat on a fresh server, two runs of each version, time to first token:

0.2.1 0.2.2
first message (3K tokens) 4.3-4.9 s 3.5-4.0 s
second turn (~570 tokens) 1.5-1.6 s 1.1-1.2 s

Later turns vary with the answers.

Also in this release

  • Output:
    • A prompt that arrives as one batch of 512 tokens or more gives bit for bit 0.2.1's output.
    • The smaller batches of a chat turn are summed in another order, so an answer can differ from 0.2.1's where two tokens were nearly tied.
    • STRIX_MMB_BF16_MIN_T=512 STRIX_MMQ_Q6K_ANY=0 in the server's environment restores 0.2.1's routing.
  • Output check:
    • The 18.6K-token equivalence probe is bit for bit 0.2.1's.
    • The long-output test reaches all 30 sections twice.
    • A 40-turn agent workload on four slots logs no warning.
    • Perplexity at 8K context is 2.6814 against 2.6811. The perplexity tool asks the output projection for 8192 rows at a time, which now run on MMQ; with STRIX_MMQ_Q6K_ANY=0 it gives 2.6811. The server asks for 1-16 rows, which MMQ already handled.

Install over 0.2.1; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 52-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/chat-ttft-20260925.json.

Requirements are unchanged:

  • Ryzen AI Max+ 395 (gfx1151)
  • Windows 11
  • a 96 GB GPU carve
  • the model files from unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ4_XS)
  • the draft head asset from 0.1.2, if you want the last 5–9% of decode