Skip to content

Releases: rulith-dev/rulith-inference

Rulith Inference 0.3.8: long conversations are summarized instead of failing, tools in several chats at once, downloads from the chat

Choose a tag to compare

@nvwaonline nvwaonline released this 30 Sep 02:59

Rulith Inference 0.3.8 keeps long conversations going by summarizing their earlier part, lets several conversations use tools at the same time, makes the download buttons in the chat work, and gives the chat the same width as the model pages.

Long conversations no longer run out of context. A conversation that grew past the model's context (256K tokens by default) failed with "Model ran out of context size", and every later message did too. Now, when the next message would take the conversation past 80% of the context, the model first writes a summary of everything so far, and the conversation carries on from the summary and the new message; the chat says when this happens and how much smaller the conversation became. The summary is written right after the last answer, while the server still has the conversation cached, so it takes about as long as a normal answer instead of re-reading the whole conversation, and the messages after it reuse the cache as before. If you edit or delete a message the summary covers, the summary is dropped. If no summary can be made, the oldest messages are left out instead, with a note telling the model so.

Tools in several conversations at once. Two conversations that used web search at the same time could leave one of them stopped after its search: it showed its sources and never answered. Switching to another chat while a search was running also stopped it. Each conversation now runs its own tool calls and answers even while you are in another chat; only a tool call waiting for your approval stops when you leave its chat, as before.

Downloads from the chat work. The download button on a code block in an answer (issue #5), a table's CSV export and the Logs page's log saved nowhere you could see. They now save into your Downloads folder, and a notice says where, with a button to show the file.

One column width. The conversation, the chat box, the new-chat page and a project's page now use the same column as the Configuration and Logs pages: up to 1280 pixels, centred.

Everything else is 0.3.7: the same model runtime, byte for byte.

Rulith Inference 0.3.7: Add documents works right after the update

Choose a tag to compare

@nvwaonline nvwaonline released this 29 Sep 23:31

Rulith Inference 0.3.7 fixes "Add documents or files" staying grey after the update from Strix Llama 0.3.5.

Documents without waiting for the model. Jan offers "Add documents or files" only to a model marked as able to use tools. 0.3.6 marked the model when it finished loading, but the model an earlier version had saved carried no mark, so after the update the menu item stayed grey, and dropped files were ignored, until a model had loaded. A document goes into the message as text, which needs nothing of the model, so the menu item and dropping files now always work; the model is also marked as soon as the app starts, which shows web search right away too.

Everything else is 0.3.6: the same model runtime, byte for byte.

Rulith Inference 0.3.6 (formerly Strix Llama): documents and web search in the chat, several conversations faster

Choose a tag to compare

@nvwaonline nvwaonline released this 29 Sep 22:03

Rulith Inference 0.3.6 is the new name of Strix Llama. It adds documents and web search to the chat, and makes several conversations at once and long conversations faster.

A new name, the same app. Strix Llama is now Rulith Inference. The old name was too close to halo-box/strix-llama.cpp, a community llama.cpp fork for the same machine that had it first. An installed Strix Llama updates in place: the same folder, settings, conversations and caches, with its Start menu and desktop shortcuts renamed. The repository moved to rulith-dev/rulith-inference, and GitHub redirects the old links.

Documents in the chat. Add documents and dropping files onto the chat did nothing until now (issue #4). They work: PDF, Word, text, Markdown and the other formats Jan's parser reads. The document's text goes into the message, so the model reads all of it; a document too long for the conversation's context is refused with a message saying so. Images can be dropped as before when the vision projector is loaded.

Web search. The model can use Jan's web search. It is off by default; the globe button in the chat box turns it on.

Faster with several conversations and long contexts. Eight conversations of about 40K tokens answering together in a 512K pool with MTP off: 75.7 to 80.8 tokens a second in total, and 74.8 to 81.6 with typical sampling settings, measured alternately with 0.3.5 in one session. One conversation at about 210K tokens: 23.2 to 24.3 tokens a second. The sparse attention now picks the blocks each token reads with one GPU launch instead of about twelve, scores each conversation's blocks only against that conversation's tokens, and reads their keys from adjacent memory instead of 1 KB apart, a spacing this machine's memory channels handle badly; the matrix products of five to eight conversations also read their inputs faster. The output is exactly the same as 0.3.5's.

Uninstalling cleans up. The model server keeps its settings and prompt cache in the app's folder, and uninstalling left them behind, often tens of GB. Ticking "Delete the application data" now removes them; without it they stay for a reinstall.

Checked before release. Identical text and probabilities to 0.3.3-0.3.5 with MTP off, with MTP on, and with a q8_0 512K pool on eight slots, and eight conversations alone and together give the same tokens as 0.3.5. Every new path was also run beside the old one, comparing each result, across eight conversations, sampling, agents, a full pool, MTP and images, with no difference. Plus the usual checks: long answers with MTP, agents with regenerates and stopped answers, requests waiting for room in a full pool, images in several conversations, the answer edges, a used slot's block keys, conversations read back from disk, and conversations moving inside the KV pool.

Known limits.

  • Documents go into the message whole; there is no search over large collections of documents (no embedding model).
  • A conversation restored from the disk cache gets its system prompt's cached values as the first conversation with the same system prompt computed them, so at temperature 0 it can occasionally answer a late token differently than if it had never left the slot.
  • A conversation with an image in it keeps the slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • Many conversations at once share 24 snapshots of the model's running state (about three each with eight), so going back far into an older part of a conversation can process more than it would alone.
  • A request waiting for room in a full KV pool receives nothing until it starts, so a client that gives up quickly on a silent request can time out first.
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.3.5: restored conversations keep all their memory; several conversations faster with MTP off

Choose a tag to compare

@nvwaonline nvwaonline released this 29 Sep 00:04

Strix Llama 0.3.5 fixes a bug that could blank part of a conversation's memory when it came back from the cache, and makes several conversations at once a little faster with MTP off.

A restored conversation keeps all of its memory. When a conversation comes back to a slot another conversation just used, from the memory cache or the disk cache, the server clears the cells the other conversation used and writes back the returning conversation's saved state. The clearing ran in the GPU's queue while the restore wrote from the CPU, and the last part of the clearing could land after the restore, blanking the returning conversation's values in the last attention layer for up to a few hundred positions. It was never another conversation's data: those positions came back empty. But the model lost some of that conversation's context in one layer, which could change its answers. It has been there since 0.1.17. The restore now waits for the clearing to finish, and a conversation's state as it left and as it came back is byte-identical. We found it while checking llama.cpp #29092, a report of recurrent state crossing between requests on this GPU. That one does not happen here: documents written in made-up vocabularies, answered one after another in the same slot, never had a word of another document in their answers.

Faster with several conversations, MTP off. Eight conversations of about 40K tokens each answering together in a 512K pool: 75.8 to 77.9 tokens a second with greedy decoding, and 73.2 to 77.7 with typical sampling settings, where the eight samples of each step are now taken in parallel instead of one after another. The sparse attention's block scoring also skips the blocks a conversation cannot see. One conversation alone runs as before, and the output is exactly the same as 0.3.4's.

Checked before release. Identical text and probabilities to 0.3.4 with MTP off, with MTP on, and with a q8_0 512K pool on eight slots; the new paths run beside the old ones in check modes over thousands of steps with no difference; a conversation's full state hashed as it leaves a slot and again when it comes back; and the usual checks: long answers with MTP, agents with regenerates and stopped answers, requests waiting for room in a full pool, images in several conversations, conversations read back from disk, the answer edges, and a 512K pool with conversations moving to disk and back.

Known limits.

  • A conversation restored from the disk cache gets its system prompt's cached values as the first conversation with the same system prompt computed them, so at temperature 0 it can occasionally answer a late token differently than if it had never left the slot.
  • A conversation with an image in it keeps the slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • Many conversations at once share 24 snapshots of the model's running state (about three each with eight), so going back far into an older part of a conversation can process more than it would alone.
  • A request waiting for room in a full KV pool receives nothing until it starts, so a client that gives up quickly on a silent request can time out first.
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.3.4: faster answers with MTP off, flat memory with many conversations

Choose a tag to compare

@nvwaonline nvwaonline released this 28 Sep 16:26

Strix Llama 0.3.4 makes answers faster with MTP off, most of all with several conversations at once and long contexts, and stops the server's memory from growing with every turn.

Faster answers with MTP off. With MTP off, a 512K-token KV pool fits beside the larger models, which makes it the better setting for many conversations at once, so this release works on the time each token costs without MTP. Three things got cheaper. The model's sparse attention first scores every block of the context to pick the ones to read; that scoring was about twenty small passes over the GPU and is now one, reading each block once. Before every token, the server worked out which parts of the KV pool each conversation may see by walking the whole pool on the CPU while the GPU waited, which cost the most with several conversations; a text conversation is one unbroken run in the pool, so it is now described from where it starts. And some per-token host work is now done once and kept. Eight conversations of about 40K tokens each answering together in a 512K pool: 58.8 to 75.7 tokens a second in total. One conversation at 110K tokens: 21.6 to 24.6 tokens a second; the slowdown as a conversation grows is less than half what it was. With MTP on, the same work goes from every pass: a long answer at 79K tokens went from 33 to 35 tokens a second. The output is exactly the same as 0.3.3's, with MTP on or off.

Memory stays flat with many conversations. The server saves snapshots of the model's running state, about 113 MB each, at points a later request may need to go back to, up to eight per conversation. With eight conversations they added about 2 GB with every turn. All conversations now share a budget of 24 snapshots (about 2.7 GB); past it, the conversation holding the most gives up its least useful one, never the one at the end of its system prompt. Eight conversations taking three turns: 8.3 GB before, 4.8 GB now, with nothing processed again (a main agent with five sub-agents did the same work as before). One conversation on its own still keeps eight.

Checked before release. Identical text and probabilities to 0.3.3 with MTP off, with MTP on, and with a q8_0 512K pool on eight slots; both new paths were also run beside the old ones, comparing every result, across eight conversations, MTP, agents and a full pool (over 20,000 comparisons, no difference). Plus the usual release checks: perplexity over 40 chunks, long answers with MTP, agents with regenerates, edits and stopped answers, rewinds at 60K and 150K characters, images in several conversations, conversations read back from disk, requests waiting for room in a full pool, the context's edges, and a 768K-cell pool with MTP.

Known limits.

  • A conversation with an image in it keeps the previous, slower path for both the scoring and the visibility work, and the first image after the model loads takes longer (its kernels load on first use).
  • Many conversations at once share the 24 snapshots (about three each with eight), so going back far into an older part of a conversation, such as an edit several turns up, can process more than before.
  • A request waiting for room in a full KV pool receives nothing until it starts, so a client that gives up quickly on a silent request can time out first.
  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.3.3: requests wait instead of failing when the KV pool is full; total throughput

Choose a tag to compare

@nvwaonline nvwaonline released this 28 Sep 08:35

Strix Llama 0.3.3 lets requests wait for room when the conversations together need more KV cache than the pool holds, instead of failing them, and shows the server's total throughput on the Logs page.

Requests wait instead of failing when the KV pool is full. With several conversations at once - agents especially - the KV pool can hold less than they need together. Until now, when a step no longer fit, every request running at that moment failed with "Context size has been exceeded", even the ones already answering, and their cached conversations were dropped. Now a request that does not fit waits in the queue until there is room: idle conversations make room first (they go to disk when Keep conversations on disk is on), and running ones finish. Running answers get room before new prompts. Only if every running conversation is stuck waiting for room does the one that arrived last give way, with an error a client can retry. Eight new 12K-token conversations sent at once to a 64K pool: 0.3.2 failed all eight, 0.3.3 answered all eight (three of them after waiting). The Logs page shows how many requests are waiting.

Throughput across all conversations. Next to Conversations on the Logs page, a line now shows how many answer tokens and prompt tokens a second the server is producing, summed over every conversation and averaged over the last 10 seconds. With several agents or chats running at once, a single conversation's speed says little; this is what the machine is doing in total. It comes from the same conversation counters the page already reads, so it costs the server nothing.

Known limits.

  • A request waiting for room receives nothing until it starts, so a client that gives up quickly on a silent request can time out first. A KV pool at least as large as the conversations you run at once avoids the wait.
  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.3.2: answers start sooner, instant regenerate, less disk writing

Choose a tag to compare

@nvwaonline nvwaonline released this 28 Sep 05:52

Strix Llama 0.3.2 makes agents and multi-turn chats start answering sooner, regenerates an answer instantly, and writes less to disk.

Checkpoints where answers begin and end, not extra passes. This model keeps part of its memory as a running state that cannot be rewound, so the server saves that state at points a later request may need to go back to. Until now it saved them by cutting each prompt into extra pieces: a few tokens before its end, and at the start of its last message. Every cut is one more pass through the model, and when several agents send requests at once, each pass waits behind their work: the last four tokens of a prompt could wait two to three seconds on their own. The server now saves the state where an answer starts and where the previous answer ended - points it passes anyway - and cuts nothing. They cover everything the cuts covered: regenerating an answer, editing the last message, a client that drops the model's reasoning from the history, a tool call the client sends back reformatted, an answer stopped halfway. A main agent with five sub-agents on eight slots: the median wait for the first token went from 5.6 to 3.0 seconds, and the time spent on prompts from 139 to 87 seconds. The longest waits, when five sub-agents send their next step at the same moment, move less (9.6 to 8.3 seconds at the 90th percentile): they are bound by how fast the prompts themselves are processed.

Regenerate is instant. The point where an answer starts keeps the probabilities of its first token, so a regenerated answer starts from them at once without processing anything (0.3.1 processed four tokens).

Requests are served in the order they arrive. Prompts waiting for the model used to be taken in slot order, so a request that came in two seconds earlier could still wait for a later one.

Less disk writing for chats that drop reasoning. With Keep conversations on disk on, the server now writes only what a conversation is known to keep: its prompts, and an answer once the next request shows the answer is kept. A client that drops the reasoning from its history re-renders every answer, and such answers are no longer written. Six conversations with thinking on taking turns in two slots: 10.8 GB written before, 8.3 GB now, with no extra processing when they come back.

Checked before release. The full release checks again: identical text and probabilities to 0.3.1 with the new checkpoints switched off (with them on, the last few prompt tokens are computed together with the rest of the prompt, as they are for everything else, which moves the first answer token's probabilities slightly), perplexity identical over 40 chunks, long answers with MTP, agents on four slots with regenerates, edits and stopped answers, rewinds at 60K and 150K characters, images in several conversations, conversations read back from disk, the context's edges, and a 768K-cell KV pool with MTP.

Known limits.

  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.3.1: a used slot's block keys, cheaper MTP drafts at depth

Choose a tag to compare

@nvwaonline nvwaonline released this 28 Sep 03:04

Strix Llama 0.3.1 fixes a quiet quality bug in long conversations and makes MTP decoding cheaper deep into a long context. Updating is recommended: the bug is in every earlier release, and the default eight conversation slots make it easy to hit.

Where it stands, at the default settings. A 95.6K-token prompt prefills at 1228 tokens a second; answers decode at 36.7 tokens a second after 86K tokens of context and 44.4 after a short question (MTP on); three and four conversations at once decode 55.4 and 62.5 tokens a second together (each step takes the same time as in 0.3.0; the sum moves with how many draft tokens get accepted).

A new conversation in a used slot reads its own beginning again. The sparse attention keeps a summary key for every block of 4 tokens. A new conversation that started in a slot another conversation had used, with a first message under about 2K tokens, kept the old conversation's summary keys for its first blocks. Once the conversation grew past 2K tokens, the model's choice of which earlier parts to attend to was scored against the wrong keys for its system prompt and first message. The answers were plausible, just not the model's: the same conversation gave different answers on a freshly started server and after another conversation in the slot. Now both give the same answer, token for token.

MTP drafts are cheaper deep into a conversation. The small draft model that proposes the next tokens read its whole context at every step. It now reuses the part of the context the model picked as relevant a few tokens earlier, as SGLang and TensorRT-LLM do for DeepSeek's sparse attention, and refreshes that choice every 32 tokens. A step of decoding (a model pass plus its drafts): 3.6% faster at 86K tokens of context, 6.5% at 212K, with the same share of drafts accepted. Below 32K tokens nothing changes. The draft also no longer computes attention for tokens it only stores, which trims a little from every step.

Checked before release. The full release checks of 0.3.0, again: text and probabilities identical to 0.3.0 with MTP off, perplexity identical chunk by chunk over 40 chunks, long answers with MTP, agents on four slots, rewinds with MTP at 60K and 150K characters, images in several conversations, conversations read back from disk, the context's edges, and a 768K-cell KV pool with MTP.

Known limits.

  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.3.0

Choose a tag to compare

@nvwaonline nvwaonline released this 27 Sep 19:43

Strix Llama 0.3.0 is the stable release of the 0.2 series: it fixes the last known way a setting could keep the model from loading, trims a little host time when several long conversations answer at once, and went through the full release checks plus a long mixed workload before release.

Where it stands, at the default settings. A 95.6K-token prompt prefills at 1217 tokens a second; answers decode at 35.8 tokens a second after 86K tokens of context and 44.9 after a short question (MTP on); three and four conversations at once decode 58.8 and 62.6 tokens a second together.

Large KV pools with MTP load. With MTP on, a KV pool above about 512K cells failed to load: the draft reserved a working buffer over the whole pool (3.9 GB at 768K). The draft now works in smaller steps when the pool is larger, so pools up to the 1M-cell maximum load. A q8_0 pool of 768K cells fits in GPU memory at four conversations; larger pools spill into shared GPU memory, which works but is slower.

Several long conversations: less host time a step. Building the attention mask for a step that serves several conversations read every cell's bookkeeping once per conversation; it now reads it once. The result is byte for byte the same.

Checked for stability. Before release it ran 150 minutes of mixed work at the default settings - agents, images in several conversations, the prompt cache, long prompts, several conversations at once - 145 workloads without a failure, its memory flat after the first half hour. Prompts longer than the context are refused with a clear message, an answer that reaches the end of the context stops cleanly, and the server keeps answering afterwards.

Known limits.

  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path, and the first image after the model loads takes longer (its kernels load on first use).
  • The same prompt can get a slightly different answer when other conversations run at the same time: batch shapes change the arithmetic in the last bits.

Strix Llama 0.2.9

Choose a tag to compare

@nvwaonline nvwaonline released this 27 Sep 14:19

A small release: sparse attention now covers every size of prompt piece, and prefill is a little faster.

Sparse attention for pieces of 33-127 tokens. In a long conversation, a piece of a prompt between 33 and 127 tokens - typical agent tool results and short follow-up messages - used dense attention over the whole conversation instead of the model's sparse attention. Such pieces now use the sparse selection like every other size.

A guard against the 0.2.7 kind of bug. If sparse attention ever reaches a kernel that would ignore its selection, the server now stops with an error instead of computing a wrong answer.

Prefill a little faster. The recurrent layers' kernel no longer spills to scratch memory, and the per-head norms run a faster kernel; the results are bit for bit the same. A 95.6K-token prompt: 1226-1231 → 1242-1245 tokens a second.

Known limits.

  • More agent sessions than Concurrent conversations, with Keep conversations on disk off: sessions push each other out and are processed again. Set Concurrent conversations to at least the number of sessions your agent runs at once.
  • A conversation with an image in it processes new text on a slower path.