Skip to content

Strix Llama 0.2.5

Choose a tag to compare

@nvwaonline nvwaonline released this 26 Sep 19:41
· 8 commits to main since this release

A redesigned app. Strix Llama now looks like the rest of Rulith's software. A new install opens on Strix Llama's own welcome screen instead of Jan's.

Rulith's design language. Light and dark themes both use:

  • neutral greys and hairline borders;
  • the system's own type: Segoe UI Variable, with Microsoft YaHei UI for Chinese;
  • one colour, Rulith green, for the model's state.

The chat, the settings and the model pages all follow it, and so does the icon, now a green tile.

The model in the sidebar. The model pages are a group of their own: Library, Configuration and Logs. At the foot of the sidebar, the model server panel shows whether the model is loaded, loading or failed. It has the one button that matters (Load, Unload or Try again), and Settings next to it.

Switch models from anywhere. Loading another model means unloading the one that runs, and the app now does both for you:

  • Pick a model from the menu in the model server panel.
  • Press Switch on its row under Model › Library.
  • Press Switch to this model on its Configuration page, which also says whether that model is the one loaded.

Load and the strip above the chat follow the model you picked. A chat that another model answered simply continues with the one now loaded.

Updates itself. From this version on, the app checks for new releases when it starts. When one is out, it offers it in a prompt, and as New version at the foot of the sidebar. The setup is installed only if its signature verifies against the key the app was built with, and the model is unloaded first. This one you still install by hand.

A welcome screen of its own. Until a model had been loaded once, a new install opened on Jan's setup screen, which offered to download Jan's own model. It now opens on Strix Llama's. The welcome screen:

  • finds the Qwen3.8 Flash Next files, or asks for their folder and links to the list of files;
  • loads the model with one button;
  • opens the chat by itself when the model is ready.

Configuration you can read. Settings are grouped in labelled sections:

  • Model
  • Generation
  • Context and memory
  • Concurrency
  • Speculative decoding
  • Image input
  • Startup
  • Advanced, collapsed

Each row says in one line what it does, and the longer explanation is behind its (i). Changing a setting no longer means save, unload and load by hand: a bar appears with Discard, Save and Save and reload.

Load when you are ready. While no model is loaded, or while one is loading, a strip above the chat says so and offers a button to load it. The app can also load the last used model when it starts: turn this on under Model › Configuration › Startup. It is off by default, so you can change the settings before anything loads.

The runtime folder no longer says 10.1. The runtime has been built with ROCm 10.2 since 0.2.3. Its folder was still named hip-rocm101, after the 10.1 it was first built with.

  • The folder is now runtime\bin\hip.
  • The setup removes the old folder when it upgrades.
  • Model › Logs shows the release: ROCm 10.2 · gfx1151.

Fixed: a crash after many images in one conversation (#2). With images in a conversation, every later batch prepares extra inputs for sparse attention. Those inputs grow with the conversation, and each time they grew, the memory they had used was not returned to Windows. 27 images of 1920×1080 in a 58K conversation used up 12 GB of RAM, and the next long text message stopped the server with unspecified launch failure.

Batches in a conversation with images are now bounded, and the buffer grows in larger steps, so it is replaced a couple of times instead of once per image. The same workload now completes with about 11 GB of RAM to spare instead of 3. Image prefill runs at the same speed. A long text message after images is about 6% slower than it would be at full batch size. Conversations without images are unchanged; their output is identical to 0.2.4's.

Conversations keep their cache past midnight. The default assistant's instructions ended with today's date, and that line sits in the system prompt at the start of every conversation. At midnight it changed, and with it the start of every prompt, so nothing cached before midnight could be reused. A 51K-token conversation continued after midnight was processed again from the start (44 s). The date line is gone from the default assistant, for new conversations and existing ones alike. An existing conversation is processed once more the first time you continue it after the update, and never again for the date. The model is no longer told the date. If you need it, add it to your assistant's instructions; it will then cost the same cache once a day.

The chat's sampling settings reach the model. Repeat penalty and Min P in the chat's parameter panel were never sent to the model. The app treated the local server like any OpenAI-compatible service, and those keys are specific to llama.cpp, so it dropped them. They are sent now.

Jan's default assistant came with a repeat penalty of 1.12, which the model therefore never ran with. It is gone from the default assistant and from existing conversations, so nothing changes for them. On this model a repeat penalty costs about a tenth of the speed: MTP's drafts are accepted less often (Chinese prose, 400 tokens, two seeds: acceptance 0.49 and 0.54 without it, 0.40 and 0.44 at 1.12; 32–34 against 29–30 tok/s). A penalty you set yourself is sent as set.

A reply cut off by an error can be continued. When the server stopped a reply part-way, the text so far was saved as if the reply had finished. That happened to every reply in progress when memory ran out. Such a reply is now saved as stopped, like one you stop yourself: it offers Continue, and continuing replaces it instead of leaving the fragment in the conversation. A model that sees fragments of its own replies in the history tends to copy them, shorter each time.

Also in this release

  • Answers cut off for lack of memory are now explained. On this machine the model, the shared context pool and every conversation in flight count against Windows commit (RAM plus page file). When it runs out, the server stops every answer being generated, and the chat used to show only an answer that ended mid-sentence. The app now says so, with the time and the number of answers, and what to change: a smaller shared context pool, fewer concurrent conversations, or a larger page file.
  • Model › Logs shows what each conversation slot is doing, as LM Studio does. It lists which slots are generating and how many tokens so far, which are reading a prompt, and which are idle with their conversation still cached, with each conversation's length. The model's API name, the one clients send as model, sits under the model's name with a copy button.
  • The model pages are wider, so log lines and paths fit. Long log lines scroll sideways unless you turn on Wrap long lines.
  • The tray menu says Open Strix Llama, and the download button Jan showed beside the sidebar toggle is gone.
  • Concurrent conversations and the Shared context pool have a section of their own. The pool used to appear only after you raised the number of conversations under Advanced. It is now always shown, and says when it takes effect.
  • Model folders are added with the system's folder picker, or by pasting a path.
  • A finished or failed load is announced wherever you are in the app.
  • The failure of the last load is shown on every model page, with a link to the log.
  • Jan's settings pages use the same layout as the model pages.
  • The accent colour picker is gone from Settings › Appearance, because the palette is fixed. Colored user message bubble now tints your messages blue.
  • Building the overlay from a fresh Jan checkout failed its type check, because the route tree did not list the model pages. The web app now type-checks after Vite writes the route tree.
  • A GPU fault in the model server now names the operation that failed and where its data lies, so a report of one is easier to trace.

Install over 0.2.4. Chats, settings and saved conversations stay, including the disk cache's. The runtime is the 53-file delta on pwilkin/llama.cpp f5daaa3 (one patch more than 0.2.4: apply_image_dense_bound), published as the strixllama branch of rulith-dev/llama.cpp. Measurements are in docs/results/issue2-image-history-20260927.json.

Requirements are unchanged:

  • Ryzen AI Max+ 395 (gfx1151)
  • Windows 11
  • a 96 GB GPU carve
  • the model files from unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ4_XS)
  • the draft head asset from 0.1.2, if you want the last 5–9% of decode

Strix Llama is made by Rulith: verifiable execution infrastructure for AI agents.