Releases: rulith-dev/strixllama
Release list
Strix Llama 0.3.0
Strix Llama 0.3.0 is the stable release of the 0.2 series: it fixes the last known way a setting could keep the model from loading, trims a little host time when several long conversations answer at once, and went through the full release checks plus a long mixed workload before release.
Where it stands, at the default settings. A 95.6K-token prompt prefills at 1217 tokens a second; answers decode at 35.8 tokens a second after 86K tokens of context and 44.9 after a short question (MTP on); three and four conversations at once decode 58.8 and 62.6 tokens a second together.
Large KV pools with MTP load. With MTP on, a KV pool above about 512K cells failed to load: the draft reserved a working buffer over the whole pool (3.9 GB at 768K). The draft now works in smaller steps when the pool is larger, so pools up to the 1M-cell maximum load. A q8_0 pool of 768K cells fits in GPU memory at four conversations; larger pools spill into shared GPU memory, which works but is slower.
Several long conversations: less host time a step. Building the attention mask for a step that serves several conversations read every cell's bookkeeping once per conversation; it now reads it once. The result is byte for byte the same.
Checked for stability. Before release it ran 150 minutes of mixed work at the default settings - agents, images in several conversations, the prompt cache, long prompts, several conversations at once - 145 workloads without a failure, its memory flat after the first half hour. Prompts longer than the context are refused with a clear message, an answer that reaches the end of the context stops cleanly, and the server keeps answering afterwards.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
- The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.
Strix Llama 0.2.9
A small release: sparse attention now covers every size of prompt piece, and prefill is a little faster.
Sparse attention for pieces of 33-127 tokens. In a long conversation, a piece of a prompt between 33 and 127 tokens - typical agent tool results and short follow-up messages - used dense attention over the whole conversation instead of the model's sparse attention. Such pieces now use the sparse selection like every other size.
A guard against the 0.2.7 kind of bug. If sparse attention ever reaches a kernel that would ignore its selection, the server now stops with an error instead of computing a wrong answer.
Prefill a little faster. The recurrent layers' kernel no longer spills to scratch memory, and the per-head norms run a faster kernel; the results are bit for bit the same. A 95.6K-token prompt: 1226-1231 → 1242-1245 tokens a second.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path.
Strix Llama 0.2.8
A fix for 0.2.7: with several long conversations loaded, a prompt that arrived while others were answering could take minutes and be answered from a wrong computation. Prompts next to long answering conversations are also faster now.
What went wrong. 0.2.7 put a new prompt and the conversations that were answering into one pass. When the conversations loaded together held more than about 262,000 tokens - easy with a larger KV pool, with the default pool only when it is nearly full - the sparse attention's fast kernel could not take such a pass, and the fallback computed the prompt's attention over every token in memory, other conversations included. A few thousand tokens then took over a minute, the other conversations stood still meanwhile, and the conversation's answer came from a wrong computation (in our test it read differently from the correct one, though still fluent). The fast kernel now takes a pass of any size.
Stored conversations are cleared once. 0.2.7 may have kept such a conversation in Keep conversations on disk, so 0.2.8 empties that store on its first start. A conversation you come back to is processed again once.
A long prompt runs in a pass of its own when that is faster. Sharing a pass costs a prompt the other conversations' contexts, so next to long conversations a prompt now runs separately; short requests - answers in progress, agent tool results next to small conversations - still share one pass. Four slots with three long conversations answering: a 49K-token prompt 793 → 1035 tokens a second, a 74K-token prompt 695 → 1006, a 2.4K-token new chat 4.8 → 3.3 s; a conversation read back from disk with a 3.9K-token tail 76.6 → 4.3 s, the others pausing for 3 s instead of 41. GPU memory is unchanged.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path.
Strix Llama 0.2.7
Several agents at once: agent sessions keep their conversations, eight stay loaded, and a long prompt no longer holds the others for many seconds at a time. Six agents on one system prompt, 0.2.6 with its four slots against 0.2.7 with its eight: tokens processed 75K → 45K, median time to the first token 10.1 → 3.4 s, the longest pause of a streaming agent 6.2 → 2.7 s.
Agent sessions keep their place. Agent tools run a session per sub-agent, all on one long system prompt, so any two sessions share most of their tokens. The server gave a new session whichever slot shared the most with it, often another agent's, and cut that agent's history, which then had to be processed again on its next step. A slot now takes a request for what it holds only when the request continues that conversation; anything else goes to a free slot and copies the shared system prompt on the GPU.
Eight conversations stay loaded. Concurrent conversations is now 8 by default, and can go up to 16. Six agents on four slots kept pushing each other out: nearly every step went back to the system prompt. With eight they stay put. Each conversation beyond the first costs about 0.4 GB of GPU memory, and a conversation in use keeps up to about 0.9 GB of snapshots in RAM. A configuration saved with an earlier version keeps its value: raise it there if you run agents.
One pass for all of them. When conversations were answering while new prompts came in, or several prompts arrived at once, the model made one pass over all its weights for each distinct prompt length: five sub-agents starting together took six passes. It is one pass now.
A long prompt no longer holds the others for long. While conversations are answering, a step now takes at most 2048 tokens of a new prompt. A 19K-token prompt used to hold a streaming conversation for 7 s at a time; now it pauses it for about 2 s at a time, and the long prompt itself takes about a fifth longer.
Update progress. The update prompt and the sidebar now show how far a download has got.
Known limits.
- More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
- A conversation with an image in it processes new text on a slower path.
Strix Llama 0.2.6
Several conversations and agents without reprocessing: conversations stay loaded, a shared system prompt is computed once, and decoding several conversations at once is a little faster.
Four conversations stay loaded. Concurrent conversations is now 4 by default (Model › Configuration › Concurrency). Each conversation keeps its place in the model, so going back to one continues where it was instead of processing it again, and up to four answer at the same time. They share one context pool, so a single conversation can still use all of it. Each conversation beyond the first costs about 0.4 GB of GPU memory. A configuration saved with an earlier version keeps its value: set it to 4 there if it says 1.
Image input with several conversations. Image input used to need Concurrent conversations at 1: an image in a second loaded conversation stopped the server. That is fixed, and both work together.
Agents: a shared system prompt is computed once. Agent tools start many sessions with the same long system prompt and tool list. A new session now takes the part it shares from a session that already computed it, copied on the GPU in tens of milliseconds, instead of computing it again. Sessions that start at the same moment wait for the first one to get past that part, then copy it. An agent workload with a 10K-token system prompt, three sessions started together and three sub-agents: 29K tokens processed instead of 73K, prompt time 172 → 60 s, the three first answers after 15–17 s instead of 36–41 s.
Edits and regenerations reuse what they can. The model can resume a conversation only at the points where it kept a snapshot of its state. Those now sit where a later message is likely to branch: where the system prompt ends, at each of your messages, and at the end of each prompt. Editing an earlier message resumes from that message instead of from the start (the second message of an 18K-token conversation: 4.3K tokens processed before, 1.1K now). Regenerating and editing the last message were already cheap and still are.
Faster with several conversations at once. When several conversations decode together, each step checked the drafted tokens of all of them and saved the model's state after every checked token, 3 MB per layer, in case a draft was rejected. It now keeps only what it needs to recompute a rejected tail. Each step at three and four conversations is 3–4% shorter. Output is identical to 0.2.5.
Known limits.
- More conversations than Concurrent conversations, with Keep conversations on disk off (the default): a conversation that had to make room is processed again when it comes back. Raise Concurrent conversations, or turn on Keep conversations on disk.
- While one conversation starts or resumes, the others pause briefly: about 0.1–0.3 s, over a second when it is read back from disk. Removing that pause is the next release's work.
Strix Llama 0.2.5
A redesigned app. Strix Llama now looks like the rest of Rulith's software. A new install opens on Strix Llama's own welcome screen instead of Jan's.
Rulith's design language. Light and dark themes both use:
- neutral greys and hairline borders;
- the system's own type: Segoe UI Variable, with Microsoft YaHei UI for Chinese;
- one colour, Rulith green, for the model's state.
The chat, the settings and the model pages all follow it, and so does the icon, now a green tile.
The model in the sidebar. The model pages are a group of their own: Library, Configuration and Logs. At the foot of the sidebar, the model server panel shows whether the model is loaded, loading or failed. It has the one button that matters (Load, Unload or Try again), and Settings next to it.
Switch models from anywhere. Loading another model means unloading the one that runs, and the app now does both for you:
- Pick a model from the menu in the model server panel.
- Press Switch on its row under Model › Library.
- Press Switch to this model on its Configuration page, which also says whether that model is the one loaded.
Load and the strip above the chat follow the model you picked. A chat that another model answered simply continues with the one now loaded.
Updates itself. From this version on, the app checks for new releases when it starts. When one is out, it offers it in a prompt, and as New version at the foot of the sidebar. The setup is installed only if its signature verifies against the key the app was built with, and the model is unloaded first. This one you still install by hand.
A welcome screen of its own. Until a model had been loaded once, a new install opened on Jan's setup screen, which offered to download Jan's own model. It now opens on Strix Llama's. The welcome screen:
- finds the Qwen3.8 Flash Next files, or asks for their folder and links to the list of files;
- loads the model with one button;
- opens the chat by itself when the model is ready.
Configuration you can read. Settings are grouped in labelled sections:
- Model
- Generation
- Context and memory
- Concurrency
- Speculative decoding
- Image input
- Startup
- Advanced, collapsed
Each row says in one line what it does, and the longer explanation is behind its (i). Changing a setting no longer means save, unload and load by hand: a bar appears with Discard, Save and Save and reload.
Load when you are ready. While no model is loaded, or while one is loading, a strip above the chat says so and offers a button to load it. The app can also load the last used model when it starts: turn this on under Model › Configuration › Startup. It is off by default, so you can change the settings before anything loads.
The runtime folder no longer says 10.1. The runtime has been built with ROCm 10.2 since 0.2.3. Its folder was still named hip-rocm101, after the 10.1 it was first built with.
- The folder is now
runtime\bin\hip. - The setup removes the old folder when it upgrades.
- Model › Logs shows the release: ROCm 10.2 · gfx1151.
Fixed: a crash after many images in one conversation (#2). With images in a conversation, every later batch prepares extra inputs for sparse attention. Those inputs grow with the conversation, and each time they grew, the memory they had used was not returned to Windows. 27 images of 1920×1080 in a 58K conversation used up 12 GB of RAM, and the next long text message stopped the server with unspecified launch failure.
Batches in a conversation with images are now bounded, and the buffer grows in larger steps, so it is replaced a couple of times instead of once per image. The same workload now completes with about 11 GB of RAM to spare instead of 3. Image prefill runs at the same speed. A long text message after images is about 6% slower than it would be at full batch size. Conversations without images are unchanged; their output is identical to 0.2.4's.
Conversations keep their cache past midnight. The default assistant's instructions ended with today's date, and that line sits in the system prompt at the start of every conversation. At midnight it changed, and with it the start of every prompt, so nothing cached before midnight could be reused. A 51K-token conversation continued after midnight was processed again from the start (44 s). The date line is gone from the default assistant, for new conversations and existing ones alike. An existing conversation is processed once more the first time you continue it after the update, and never again for the date. The model is no longer told the date. If you need it, add it to your assistant's instructions; it will then cost the same cache once a day.
The chat's sampling settings reach the model. Repeat penalty and Min P in the chat's parameter panel were never sent to the model. The app treated the local server like any OpenAI-compatible service, and those keys are specific to llama.cpp, so it dropped them. They are sent now.
Jan's default assistant came with a repeat penalty of 1.12, which the model therefore never ran with. It is gone from the default assistant and from existing conversations, so nothing changes for them. On this model a repeat penalty costs about a tenth of the speed: MTP's drafts are accepted less often (Chinese prose, 400 tokens, two seeds: acceptance 0.49 and 0.54 without it, 0.40 and 0.44 at 1.12; 32–34 against 29–30 tok/s). A penalty you set yourself is sent as set.
A reply cut off by an error can be continued. When the server stopped a reply part-way, the text so far was saved as if the reply had finished. That happened to every reply in progress when memory ran out. Such a reply is now saved as stopped, like one you stop yourself: it offers Continue, and continuing replaces it instead of leaving the fragment in the conversation. A model that sees fragments of its own replies in the history tends to copy them, shorter each time.
Also in this release
- Answers cut off for lack of memory are now explained. On this machine the model, the shared context pool and every conversation in flight count against Windows commit (RAM plus page file). When it runs out, the server stops every answer being generated, and the chat used to show only an answer that ended mid-sentence. The app now says so, with the time and the number of answers, and what to change: a smaller shared context pool, fewer concurrent conversations, or a larger page file.
- Model › Logs shows what each conversation slot is doing, as LM Studio does. It lists which slots are generating and how many tokens so far, which are reading a prompt, and which are idle with their conversation still cached, with each conversation's length. The model's API name, the one clients send as
model, sits under the model's name with a copy button. - The model pages are wider, so log lines and paths fit. Long log lines scroll sideways unless you turn on Wrap long lines.
- The tray menu says Open Strix Llama, and the download button Jan showed beside the sidebar toggle is gone.
- Concurrent conversations and the Shared context pool have a section of their own. The pool used to appear only after you raised the number of conversations under Advanced. It is now always shown, and says when it takes effect.
- Model folders are added with the system's folder picker, or by pasting a path.
- A finished or failed load is announced wherever you are in the app.
- The failure of the last load is shown on every model page, with a link to the log.
- Jan's settings pages use the same layout as the model pages.
- The accent colour picker is gone from Settings › Appearance, because the palette is fixed. Colored user message bubble now tints your messages blue.
- Building the overlay from a fresh Jan checkout failed its type check, because the route tree did not list the model pages. The web app now type-checks after Vite writes the route tree.
- A GPU fault in the model server now names the operation that failed and where its data lies, so a report of one is easier to trace.
Install over 0.2.4. Chats, settings and saved conversations stay, including the disk cache's. The runtime is the 53-file delta on pwilkin/llama.cpp f5daaa3 (one patch more than 0.2.4: apply_image_dense_bound), published as the strixllama branch of rulith-dev/llama.cpp. Measurements are in docs/results/issue2-image-history-20260927.json.
Requirements are unchanged:
- Ryzen AI Max+ 395 (gfx1151)
- Windows 11
- a 96 GB GPU carve
- the model files from
unsloth/Qwen3.8-Flash-Next-GGUF(UD-IQ4_XS) - the draft head asset from 0.1.2, if you want the last 5–9% of decode
Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.
Strix Llama 0.2.4
A small release for several conversations at once: a verify step of three or four conversations is 3-5% shorter, which gives 2-4.5% more tokens. It also corrects the numbers 0.2.3 published.
Three more small products off few-workgroup kernels. With speculative decoding, three conversations drafting two tokens each verify nine tokens a step. After 0.2.3, three products of that step still ran on kernels that gave the whole GPU only a few workgroups:
- The input of the gated delta net layers' convolution, which appends each conversation's new tokens to its conv state. Below 32 tokens it took a generic copy kernel with one block per channel and conversation, about 30,000 small blocks at three conversations. It now takes the tiled transpose from 2 tokens.
- The sparse-attention indexer's BF16 key projection, from 9 tokens on. It ran on one 128-row tile: one workgroup.
- Quantized weights with at most 1024 output rows, from 9 tokens on: the hyper-connection down projection (96 times a step), the attention keys and values, the shared expert. Three to five tiles each.
The last two now run the vector kernel over chunks of 8 tokens. By the server's own timing over ~1,900 steps per build, a nine-token verify step takes 103.0 ms instead of 106.4, a twelve-token one 123.5 instead of 129.4.
0.2.3 against 0.2.4, alternating, three passes of three rounds. Each conversation has ~4K tokens of context and generates 512 tokens with the server's default sampling, with a new seed every round (the same seeds for both builds). How many drafts are accepted depends on the text a round happens to sample (54-86% here), and that moves tok/s more than this change does, so the table compares the builds at the same acceptance:
| 0.2.3 | 0.2.4 | |
|---|---|---|
| three conversations, at 68% acceptance | 53.8 tok/s summed | 55.0 (+2.2%) |
| four conversations, at 67% | 58.1 | 60.7 (+4.5%) |
| one conversation | 37.5 | 37.6 (the same output bit for bit) |
A correction to 0.2.3's numbers. The 0.2.3 notes said three conversations got 17% more tokens than with 0.2.2, and four 9% more. That comparison fixed each conversation's sampling seed. A server that computes every step the same way then samples the same text in every round, and the acceptance is that text's: 0.2.2 sampled nearly the same text every round (66-67%), 0.2.3 did not (65-72%), so its average was partly the luck of the texts it drew. At the same acceptance, 0.2.3 got 13% more than 0.2.2 at three conversations, and 8% more at four.
Also in this release
- With fixed seeds, three conversations now repeat their text from run to run: the same texts in four runs out of four, where 0.2.3 gave three different outcomes in four. Up to 0.2.3 these products changed kernel at 9 tokens, so a token's result could depend on how many tokens its step happened to carry. Four conversations can still vary.
- Output:
- The decode of one or two conversations, and any batch of more than 32 tokens, give bit for bit 0.2.3's output.
- A batch of 9-32 tokens is summed in another order: the verify step of three or four conversations, a short chat turn, the last few tokens of a prompt. An answer can differ from 0.2.3's where two tokens were nearly tied.
STRIX_BF16_VEC_CHUNK_MAX=0 STRIX_Q_VEC_CHUNK_MAX=0in the server's environment restores 0.2.3's routing.
- Output check:
- The 18.6K-token equivalence probe is bit for bit 0.2.3's.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity at 8K context is 2.6814, as in 0.2.3. With 16-token batches, where every batch takes the new paths, it is 4.0732 against 4.0870 with 0.2.3's routing.
Install over 0.2.3; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 53-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/verify-small-products-20260926.json.
Requirements are unchanged:
- Ryzen AI Max+ 395 (gfx1151)
- Windows 11
- a 96 GB GPU carve
- the model files from
unsloth/Qwen3.8-Flash-Next-GGUF(UD-IQ4_XS) - the draft head asset from 0.1.2, if you want the last 5–9% of decode
Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.
Strix Llama 0.2.3
Several conversations at once are faster: three conversations decoding together get 17% more tokens, four 9% more.
A cheaper verify step for several conversations. With speculative decoding, each step verifies the drafted tokens of every conversation that is generating: three conversations drafting two tokens each verify nine tokens a step. At that size two of the model's products fell off their fast kernel:
- The MoE router, from 9 tokens on. It ran on tiles of 128 rows, four workgroups for the whole GPU, about 250 µs a product.
- The routed experts, from 5 tokens on. They ran on a tiled kernel that stages a tile per expert for the one or two tokens each expert gets.
Both now run the fast kernel over chunks of the batch. A nine-token verify step takes 103 ms instead of 120.
Drafts for four conversations. With the cheaper step, four conversations drafting two tokens each beat not drafting: 62.6 against 55.7 tok/s summed at ~4K tokens of context each, 56.1 against 50.6 at ~20K. The draft is now 3 tokens for one conversation, 2 for two to four, and none from five on. Six and eight conversations still do better without drafts.
ROCm 10.2. The runtime is built against TheRock ROCm 10.2.0a20260925. On the same source it gives the same perplexity as 10.1 and the same output bit for bit on the 18.6K-token equivalence probe. Prefill and single-conversation decode are unchanged; three or four conversations get 3-4% more.
0.2.2 against 0.2.3, alternating, three passes of two rounds each. Each conversation has ~4K tokens of context and generates 512 tokens with the server's default sampling. Means of six rounds:
| 0.2.2 | 0.2.3 | |
|---|---|---|
| one conversation | 38.0 tok/s | 37.5 |
| three conversations | 47.7 tok/s summed (1.26× one) | 55.7 (1.47×) |
| four conversations | 55.4 (1.46×) | 60.4 (1.59×) |
One conversation runs the same code as before; the 0.5 tok/s is one slower pass at the same 63% draft acceptance. With three and four conversations the acceptance depends on the sampled text, so single rounds spread by up to 5%.
Also in this release
- The Strix Llama settings page has a one-line credit at the bottom.
- Output:
- One conversation's decode, and any batch of more than 32 tokens, give bit for bit 0.2.2's output.
- A batch of 5-32 tokens is summed in another order: the verify step of several conversations, a short chat turn, the last few tokens of a prompt. An answer can differ from 0.2.2's where two tokens were nearly tied.
STRIX_F32_VEC_CHUNK_MAX=0 STRIX_MOE_VEC_CHUNK=0in the server's environment restores 0.2.2's routing.
- Output check:
- The 18.6K-token equivalence probe is bit for bit 0.2.2's.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity at 8K context is 2.6814, as in 0.2.2.
Install over 0.2.2; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 52-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/multi-stream-20260926.json.
Requirements are unchanged:
- Ryzen AI Max+ 395 (gfx1151)
- Windows 11
- a 96 GB GPU carve
- the model files from
unsloth/Qwen3.8-Flash-Next-GGUF(UD-IQ4_XS) - the draft head asset from 0.1.2, if you want the last 5–9% of decode
Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.
Strix Llama 0.2.2
Chats get to their first token sooner.
No first-use stalls in a chat turn. hipBLAS loads each GEMM kernel from disk the first time it meets a shape. Two of the model's products still went through it:
- The sparse-attention indexer's BF16 projections, in any batch of 17 to 511 tokens. That is the new part of most chat turns, and it cost about 0.4 s the first time.
- The MTP draft head's Q6_K projection, in the draft's first prompt batch. That cost about 0.2 s.
Both now stay on this fork's own kernels.
A five-turn chat on a fresh server, two runs of each version, time to first token:
| 0.2.1 | 0.2.2 | |
|---|---|---|
| first message (3K tokens) | 4.3-4.9 s | 3.5-4.0 s |
| second turn (~570 tokens) | 1.5-1.6 s | 1.1-1.2 s |
Later turns vary with the answers.
Also in this release
- Output:
- A prompt that arrives as one batch of 512 tokens or more gives bit for bit 0.2.1's output.
- The smaller batches of a chat turn are summed in another order, so an answer can differ from 0.2.1's where two tokens were nearly tied.
STRIX_MMB_BF16_MIN_T=512 STRIX_MMQ_Q6K_ANY=0in the server's environment restores 0.2.1's routing.
- Output check:
- The 18.6K-token equivalence probe is bit for bit 0.2.1's.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity at 8K context is 2.6814 against 2.6811. The perplexity tool asks the output projection for 8192 rows at a time, which now run on MMQ; with
STRIX_MMQ_Q6K_ANY=0it gives 2.6811. The server asks for 1-16 rows, which MMQ already handled.
Install over 0.2.1; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 52-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/chat-ttft-20260925.json.
Requirements are unchanged:
- Ryzen AI Max+ 395 (gfx1151)
- Windows 11
- a 96 GB GPU carve
- the model files from
unsloth/Qwen3.8-Flash-Next-GGUF(UD-IQ4_XS) - the draft head asset from 0.1.2, if you want the last 5–9% of decode
Strix Llama 0.2.1
Several conversations at once: with speculative decoding on, three conversations decoding together are about a third faster.
Drafts of one length. With several conversations generating, the MTP drafter stopped each conversation's draft at its own first unconfident token, so one step could draft 2, 2 and 1 tokens. This model's hybrid memory processes a batch in pieces with the same number of tokens for every conversation. An uneven verify therefore ran as two or three passes of the whole model, each in a shape the GPU graph cache had not seen: 200 to 440 ms, where an even one takes about 120. Three conversations at once got 34 tok/s summed with drafting on, fewer than the 46 they got with it off.
Now the drafts of one step all have the same length. A conversation that turns unconfident keeps drafting while another is still confident, and all drafts are cut to the longest confident run.
| before | now | |
|---|---|---|
| three conversations, ~4K tokens each | 34.3 tok/s summed | 45.7 |
| three conversations, ~20K tokens each | 39.6 | 45.7 |
| two conversations | swung between 26 and 41 | 39-42 |
One conversation is unchanged, and so are four or more, which do not draft.
No 270 ms stall when several conversations start decoding together. For several conversations, sparse attention computes which conversation each block of the cache belongs to. That product went to hipBLAS, which loads a kernel from disk the first time it meets a shape: 272 ms in the first step three conversations decoded together, and again as their contexts grew. It now has its own small kernel (0.02 ms, same result).
Also in this release
- Corrected: 0.2.0's notes gave 0.1.17's perplexity as 2.6880, but that was measured on a development build. On the 0.1.17 release it is 2.6841, against 0.2.0's 2.6811. On the older test (English and code, 4K context) it is 2.4741 against 2.4730. The accuracy is the same.
- Where the time of a step with several conversations goes, and why gufo's multi-user table is not comparable: every user there is sent the same prompt, so all of them route to the same experts and the weights are read once. See
docs/results.md. - Output check:
- One conversation's output is bit for bit 0.2.0's.
- With q8_0, a conversation processed over another's freed cells hashes the same as alone after every prompt batch.
- The same 85K-token prompt sent twice gives the same 400 tokens as 0.2.0.
- Two long conversations (173K and 155K tokens) read back from disk answer exactly as when both stay resident.
- The image stress test passes with the disk tier on.
- A 137K-token conversation with short prompts in between logs no assert.
- A server that failed an allocation at a checkpoint answers the retry exactly as one that never failed.
- The long-output test reaches all 30 sections twice.
- A 40-turn agent workload on four slots logs no warning.
- Perplexity is unchanged.
Install over 0.2.0; chats, settings and saved conversations stay, including the disk cache's. The runtime is the 51-file delta on pwilkin/llama.cpp f5daaa3, published as the strixllama branch of rulith-dev/llama.cpp. Measurements: docs/results/multi-stream-20260925.json.
Requirements are unchanged: Ryzen AI Max+ 395 (gfx1151), Windows 11, a 96 GB GPU carve, the model files from unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ4_XS), and the draft head asset from 0.1.2 if you want the last 5–9% of decode.