Releases: Niko1221/Strata
Release list
Strata v0.1.24
Long prompts read faster: the sparse attention's selection runs on tensor cores.
For every token of a prompt, Strata picks which earlier parts of the context its attention reads (the QSA selection). At long context that choice was 18% of the prompt time. Now:
- the scores it chooses by are computed on tensor cores, at FP32-level accuracy (3xTF32);
- the top-k choice keeps each query's scores in registers - it picks exactly the same parts as before (also while writing answers).
RTX 5070, Q2_0: a 128K-token prompt 1,608 → 1,843 tokens/s (time to first token 82.7 s → 72.3 s), a 32K prompt 1,742 → 1,797 tokens/s. The selection step itself is 3.3-3.9x faster.
Checked before the release: needle tests 15/15 (8K, 16K, 32K at five depths); with the old selection switched on (STRATA_SELECT_OLD=1 STRATA_TOPK_OLD=1) the output is byte-identical to 0.1.23 on Q2_0 and the Coder, including the prompt path's internal state at 4K and 20K; a live server test of every endpoint and image requests passes; builds and runs the same on Linux. Answers can differ from 0.1.23 in the last bits of a score (another summation order), the same kind of change as 0.1.22's attention kernel.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.24.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.23
Faster decoding, llama.cpp-compatible server endpoints, and fixes for images, 8 GB cards and multi-GPU.
Faster answers: the verify window's per-token kernels (router, norms, RoPE, the shared-expert cast, the combine) run once for the whole window instead of once per token - bit-identical output (#109, thanks @q8atnight). On an RTX 5080 + RTX 3090 decoding went 81 → 89 tokens/s on a story and 105 → 118 on code.
Server (for other apps):
GET /models,/v1/models,/propsand/slotsin llama.cpp's format, so clients like Pi find the model, its context and its image support (#122, thanks @hendrikp; fixes #117).- Replies carry
usage.prompt_tokens_details.cached_tokens(how much of the prompt the conversation cache reused) and llama.cpp'stimings(prompt and answer speed, draft counts), andGET /v1/statusgives a small summary (#107, thanks @architectds; the draft counts from #83, thanks @mikicvi). - Image parts in tool and assistant messages are kept (#113, thanks @Yunado).
- The web app works behind a reverse proxy under a path prefix (#82, thanks @mikicvi); the Monitor's live speed is a rate over the last seconds (#104, thanks @xyzzing).
Fixes:
- Image questions honor the sampling settings (temperature, top_p, top_k, penalties, seed). They always decoded greedily before, which could make the model loop on pictures (#75, thanks @jleaders).
- 8 GB cards with no room left for the expert cache start again: the prompt is read in smaller chunks, and if even those don't fit, it says what frees VRAM (#85).
- Two GPUs with an 8K context (fp16 KV): the second card's prompt attention failed to launch; fixed.
- Setup: a size that needs more RAM than the PC has is now a warning and a question instead of a stop (#125, thanks @FatheredPuma81).
- Windows: when the model loads very slowly because Strata was started by Task Scheduler, a hint says how to fix the task (#126, thanks @Zauberio; see "Running it at startup" in docs/DETAILS.md).
- Loading: the experts and the draft layer are read in large blocks on Windows (#89, thanks @dannychirkov); a GGUF with a duplicate tensor name is refused with its name (#91), and a missing or short model shard is named (#92, thanks @Avicennasis); RTX 30-class cards are accepted by the device check (#81, thanks @mikicvi); a CPU-only configure with tests works again (#93).
--dump-routingalso records the live server's routing (#118, thanks @maedoc).
Checked before the release: byte-identical to 0.1.22 on all five published quants (Windows), on Linux, and on one and two GPUs (RTX 5080 + RTX 3090), plus a live server test of every endpoint above and image requests.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.23.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.22
Faster prompts, and several GPUs without typing any flags.
Prompts are read faster (RTX 5070, Q2_0): a 32K-token prompt 1,213 → 1,646 tokens/s, a 128K prompt with KV streaming 1,144 → 1,596 tokens/s (time to first token 116 s → 83 s). Where it came from:
- the sparse attention over a prompt runs on tensor cores (4x faster at FP32-level accuracy; needle tests pass at every depth);
- the n-gram table rows are read from disk by several threads, on Windows and Linux (on Linux the reads were one at a time: a 16K prompt read 3x faster in our test);
- the selection, indexer, short-conv and GDN kernels of the prompt path, the embedding gather, the draft layer's prompt pass and the expert streaming of the IQ quants no longer wait on the host (these are bit-identical: with
STRATA_PROMPT_ATTN_OLD=1every output is byte-identical to 0.1.21); - the expert-cache slots the prompt path borrows are given back in one batch.
Output speed is unchanged.
Two or three GPUs (#112, #77, #95, #36; thanks @megakilo, @sal7g86, @pdsmike):
START-HERE.bat/./setup.shlists your NVIDIA cards, says which ones Strata can use (and why not, e.g. older than the RTX 30 series), and asks which to use, with the two best together recommended.- A model you installed on one card asks once, at its next start, whether to use both from now on.
--gpus 0,1,--gpus allor--gpu 0,1work at any start and are remembered;--gpu Nruns one card for that start.- Fixed: 0.1.21's
--setup --gpus 0,1stopped at the end with "there is no GPU [0, 1]";--gpuson an installed model started on one card; an update unzipped next to a split install stopped at step 1. - Images, the experimental speed projection (control vectors), KV streaming and mid-prompt checkpoints now work across cards;
--calibratemeasures the split; the Monitor shows every card (total VRAM, power and load, each card's own reading underneath). - The engine is compiled for every chosen card's generation when you build it yourself.
- Tested on an RTX 5080 + RTX 3090 (+ a 2080 Ti, which is refused with its reason). Linux uses the same setup code, but we could not test several GPUs on Linux; reports welcome.
Also:
- Setup compiles with Visual Studio 2022 even when Visual Studio 2026 is installed too (CUDA 13 refuses 2026).
- The streaming detokenizer no longer slows down as a reply grows (it cost 4 ms per token after 16K tokens).
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.22.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.21
One model across two or three NVIDIA cards (experimental).
- Layer split across GPUs:
- Setup:
START-HERE.bat --setup --gpus 0,2(Linux:./setup.sh --setup --gpus 0,2), or"gpu": [0, 2]in a run config. - Each card runs a range of the model's layers and keeps the experts of those layers in its own VRAM; the last card also runs the output head and the draft layer.
- Prompts flow through the cards in a pipeline: the next card reads chunk c while the first reads chunk c+1.
- Conversation checkpoints, adaptive expert swaps and the PCIe share work per card.
- No NVLink or peer-to-peer access needed: cards in x4 slots work.
- Setup:
- Measured on an RTX 5080 + RTX 3090 (Ryzen 9 9950X3D) with the Coder, 32K context:
- prompts 18-20% faster than the 5080 alone: 2,357 vs 1,970 tokens/s at 28K tokens;
- decoding on par (84 / 110 tokens/s, story / code), because both caches then hold ~99% of the routed experts and the per-layer GPU time decides.
- Placement:
--layer-split auto(the default with several cards) places the split from each card's free VRAM and speed. It picked a split within 0-12% of the best one measured.- List the fastest card first, and leave out a much slower one: an RTX 2080 Ti as a third card made the pair slower.
- Details, limits and all measurements: docs/MULTI_GPU.md and
bench/results/2026-09-29-layer-split/.
- Also included:
- The experimental helper-GPU expert caches from #16 by @Daerdaal (
--expert-cache-remote, docs/SECOND_GPU.md): the first multi-GPU groundwork. - On a PCIe link below 4 GB/s (x1 risers), no missing expert is read over PCIe any more; the CPU computes them all (it is faster there).
-DSTRATA_EXPERIMENTAL_SM75=ONbuilds the engine for an RTX 20-series card (Turing), meant as an extra layer-split card. Not supported otherwise.
- The experimental helper-GPU expert caches from #16 by @Daerdaal (
- One GPU is unchanged:
- Tested against 0.1.20 on Windows with every model size (Coder, Q2_0, IQ3_XXS, IQ3_S, Swift IQ2_XS) and on Linux (Q2_0): the same answers, byte for byte.
- Same speed.
Thanks to the two community members who lent their PCs for the multi-GPU testing.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.21 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.20
Faster new chats for agent clients, a hit-rate column in the Monitor, a PCIe-aware default, and community fixes.
- New chats reuse the system prompt (#62 by @j-luwierski, #65 by @code-martin):
- A prompt read from the start is now also checkpointed at the end of its system prompt (with the tool list), when that is 2,048 tokens or more.
- That checkpoint is kept while the others rotate by least recent use, so the next new chat of the same client reads only what comes after it.
- Measured: a 3,883-token system prompt is reused whole by a new chat, also after more turns than the cache has slots.
--prompt-cache-root 0turns it off. See the details.
- Monitor: expert-cache hit rate per request (#69 by @code-martin): the share of the experts looked up while writing the answer that were already in VRAM.
- A default PCIe share that fits your link (#44 by @pipeob0):
- The engine measures the PCIe bandwidth at start. On a narrower link (e.g. an x8 card) it lowers the share of the missing experts that the GPU reads over PCIe.
- x16 links keep the measured default.
--calibratestill measures it outright.
- OrcaRouter's Uncensored IQ3_XXS can be packed with
iq_pack.py --compat-bf16(#67 by @acrogenesis); see docs/ORCA.md. It is a manual setup, not an installer option. The official models pack exactly as before. - Fixed:
- The parity tool fails the run on a kernel mismatch instead of only printing it (#64 by @pipeob0).
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.20 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.19
Penalties that really apply, tuning for your PC, tools from MCP servers, and a fix for starting with a small page file.
- Fixed: repetition, presence and frequency penalties were mostly not applied (#53).
- What went wrong: speculative decoding checks several guessed tokens at once, but only the first of each batch got the request's penalty history. The others were judged against memory that was never filled in.
- What it meant: with ~3 tokens per step, most tokens ignored the penalties, and
penalty_last_nnear 4096 could crash the engine. - Now every checked token gets its own history, exactly as if the engine decoded one token at a time.
- The cost: requests with penalties are 1-11% slower, because more draft guesses are rejected. Requests without penalties give the same answers at the same speed.
- Found by an engineering review of the engine. Tested on Q2_0, IQ2_XS (Swift), IQ3_XXS, IQ3_S and the Coder.
- Sampler:
top_k= 0 or above 64 now means "the widest list the engine keeps" (64), instead of an unwritten token (engine) or a silent 20 (server);- token ids outside the vocabulary in a history are ignored.
strata.exeon the command line samples again: with--temperature/--seedthe speculative loop used to decode greedily. It also prints the sampling it uses.- The model starts with a small Windows page file (#60, thanks @shanoor).
- Why it failed: memory on the graphics card also needs room in Windows' commit (RAM plus page file). With the page file off or tiny, the expert cache could not be allocated even though VRAM was free, and the engine stopped.
- Now it retries with a smaller cache and says why.
- Setup warns when the page file is under 4 GB. The fix: set it to "System managed".
START-HERE.bat --calibrate: tuning for your PC.- It measures three engine settings that depend on the machine: the PCIe share, the draft depth and the CPU threads.
- It keeps a setting only when it is more than 3% faster, and remembers the result per PC and model.
- Setup offers it after an install. On the RTX 5070 test PC it made the Coder 7.6% faster.
- Tools from MCP servers in the chat page.
- List servers in the run config under
mcp_servers; Claude Desktop'smcpServersblock can be pasted as it is. - The model can then call their tools in the web app's chat: files, search, GitHub and so on. Each call shows as a card.
- API clients are not affected. See the details.
- List servers in the run config under
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.19 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.18
Clearer engine updates, a SETUP.bat shortcut, and a sampler guard.
- Updating right after a release no longer keeps the old engine without saying why (#58, thanks @XeonG). New files on
maincould reach you a few minutes before the release's engine finished uploading; setup then kept the installed engine with only a short warning. Releases are now published together with the engine, and if it ever happens again setup says the new engine isn't published yet and to run it again in a few minutes. SETUP.bat: double-click it to change settings or install another model (the same asSTART-HERE.bat --setup; on Linux./setup.sh --setup).- Sampler: with the penalty window set to 0, the sampling kernel no longer touches its (then empty) penalty buffer when a caller still passes a token history (#59, thanks @j-luwierski). The server never does this, so answers are unchanged; it's a guard for other callers of the kernel.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat. An existing 0.1.17 engine keeps working (nothing in this release needs the new one); a new install gets 0.1.18. On Linux, run ./setup.sh.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.17
Claude Code works with Strata, sampling matches llama.cpp, and you can choose the GPU.
- Claude Code can use Strata (#55, #56, thanks @Nicolas0315): the server accepts
/v1/messages?beta=true, and a system message in the middle of a conversation (Claude Code sends its hook context that way; some OpenAI clients send late "developer" messages) no longer breaks the chat template. SetANTHROPIC_BASE_URL=http://127.0.0.1:8080andANTHROPIC_MODELto a Claude model name it knows (Strata ignores the name) - see the details. - Fixed: sampling applied the repetition, frequency and presence penalties twice (#53, thanks @hendrikp). With
presence_penalty1.5 at temperature 0.7, a repeated word was about 4x less likely than it should be. The order is now llama.cpp's: penalties once, then top_k, top_p, min_p, temperature. Greedy answers (temperature 0, the default) are unchanged. - Choose the GPU on a PC with more than one (#51, thanks @ivancheg8): setup uses the one with the most VRAM, or
START-HERE.bat --gpu 1(numbered asnvidia-smishows them); the Monitor shows that card. tools/needle_bench.pychecks long-context recall through a running server: it hides a code word in a long text and asks for it (#33).
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.17 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.16
New: the Coder model. No more re-downloads after an update. Strata is now open source (MIT).
- The GSQ-RCO Coder (ISTA-DASLab), thanks to @pjgmobile's PR #54: Qwen3.8-Flash-Next with half of its experts removed - the ones code, tool use and images need are kept (91% of the full model's SWE-bench Verified, 99% of LiveCodeBench, by its authors). It needs only 32 GB of RAM, runs 262K context on 64 GB, and on a 12 GB card reads prompts at 599-1,350 tokens/s and writes 43-53 tokens/s (1K-262K). Images work. Install it with
START-HERE.bat --setup --family coder; with the original model installed, only its 29.6 GB shard 1 is downloaded (the rest is the same file). - Updating no longer downloads the model again. The model files now live in a
Strata-datafolder next to your Strata folder, so a new copy (an update unzipped elsewhere) finds them and sets itself up the same way. Existing installs are moved there at their next start - instantly on the same drive, and their old folder keeps working. - Clearer start. Before and while the model loads, the window says that your PC can be slow or stop responding for 1-3 minutes (the model is loaded into RAM and locked for the GPU), and what it is doing, instead of a silent wait.
- Open source: Strata is now under the MIT License. Thank you to everyone who contributed - @pipeob0, @pjgmobile, @code-martin, @coolio986, @Mirtraxxx, @Vistawizard, @j-luwierski and everyone who reported issues.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.16 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.
Strata v0.1.15
Fixed: IQ3_XXS stopping on long prompts, and new speed tables.
- Fixed: IQ3_XXS could stop on prompts of 64K tokens and more on a 12 GB card, with
out of memory: cudaFuncSetAttributein the log. CUDA loads a kernel's code the first time it is used, and in the middle of a long prompt there was no VRAM left for it. The engine now loads all of its kernels when it starts, before the expert cache takes the free VRAM (~30 MB, about 20 fewer experts on the GPU). Other models and shorter prompts were not affected. - New speed tables in the README and details, measured with the current engine and settings: prompts are read 1.2x (1K) to 2.4x (32K-128K) faster than the old tables showed - Q2_0 reads a 32K prompt at 1,308 tokens/s (was 571), IQ3_S at 1,070 (was 397). Raw numbers:
bench/results/2026-09-28-speed-0114.
Updating: get the latest files (git pull or download the repo again), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.15 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.