Skip to content

Rulith Inference 0.3.6 (formerly Strix Llama): documents and web search in the chat, several conversations faster

Choose a tag to compare

@nvwaonline nvwaonline released this 29 Sep 22:03
· 7 commits to main since this release

Rulith Inference 0.3.6 is the new name of Strix Llama. It adds documents and web search to the chat, and makes several conversations at once and long conversations faster.

A new name, the same app. Strix Llama is now Rulith Inference. The old name was too close to halo-box/strix-llama.cpp, a community llama.cpp fork for the same machine that had it first. An installed Strix Llama updates in place: the same folder, settings, conversations and caches, with its Start menu and desktop shortcuts renamed. The repository moved to rulith-dev/rulith-inference, and GitHub redirects the old links.

Documents in the chat. Add documents and dropping files onto the chat did nothing until now (issue #4). They work: PDF, Word, text, Markdown and the other formats Jan's parser reads. The document's text goes into the message, so the model reads all of it; a document too long for the conversation's context is refused with a message saying so. Images can be dropped as before when the vision projector is loaded.

Web search. The model can use Jan's web search. It is off by default; the globe button in the chat box turns it on.

Faster with several conversations and long contexts. Eight conversations of about 40K tokens answering together in a 512K pool with MTP off: 75.7 to 80.8 tokens a second in total, and 74.8 to 81.6 with typical sampling settings, measured alternately with 0.3.5 in one session. One conversation at about 210K tokens: 23.2 to 24.3 tokens a second. The sparse attention now picks the blocks each token reads with one GPU launch instead of about twelve, scores each conversation's blocks only against that conversation's tokens, and reads their keys from adjacent memory instead of 1 KB apart, a spacing this machine's memory channels handle badly; the matrix products of five to eight conversations also read their inputs faster. The output is exactly the same as 0.3.5's.

Uninstalling cleans up. The model server keeps its settings and prompt cache in the app's folder, and uninstalling left them behind, often tens of GB. Ticking "Delete the application data" now removes them; without it they stay for a reinstall.

Checked before release. Identical text and probabilities to 0.3.3-0.3.5 with MTP off, with MTP on, and with a q8_0 512K pool on eight slots, and eight conversations alone and together give the same tokens as 0.3.5. Every new path was also run beside the old one, comparing each result, across eight conversations, sampling, agents, a full pool, MTP and images, with no difference. Plus the usual checks: long answers with MTP, agents with regenerates and stopped answers, requests waiting for room in a full pool, images in several conversations, the answer edges, a used slot's block keys, conversations read back from disk, and conversations moving inside the KV pool.

Known limits.

  • Documents go into the message whole; there is no search over large collections of documents (no embedding model).
  • A conversation restored from the disk cache gets its system prompt's cached values as the first conversation with the same system prompt computed them, so at temperature 0 it can occasionally answer a late token differently than if it had never left the slot.
  • A conversation with an image in it keeps the slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • Many conversations at once share 24 snapshots of the model's running state (about three each with eight), so going back far into an older part of a conversation can process more than it would alone.
  • A request waiting for room in a full KV pool receives nothing until it starts, so a client that gives up quickly on a silent request can time out first.
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.