Strata v0.1.30
Short prompts up to 28% faster (the same answers), multi-GPU and AMD RDNA4 improvements, a server that can give the GPU back, and several opt-in features.
Faster short prompts. Prompts of 1K to 4K tokens read all experts from the stream from 1,024-token chunks
(0.1.29: 2,048): +17-28% prompt speed on this RTX 5070, with the same output (suggested in the issues).
STRATA_PREFILL_STREAM_MIN overrides the threshold.
Multi-GPU (#216, @gopinath87607). In a layer split:
- Each card now keeps only its own layers' session state, which saved 0.93 GiB on a later card in the author's
4-GPU test. - Each card lends the prompt path the tail of its own expert cache, instead of every card after the first
holding back 1 GiB for the whole session.--prefill autocan then pick larger chunks (8,192 tokens instead of
2,048 there). - The Monitor's expert cache counts every card (it showed only the first).
- Snapshots and checkpoints follow the split: a card saves and restores only its own layers.
Single-GPU runs are unchanged (byte-identical).
AMD RDNA4 (#178, @doplxyz). The RX 9070 / 9070 XT and the Radeon AI PRO R9700 (gfx1201) are supported beside
the RX 7900 series: setup picks ROCm and the settings per card, and the engine checks at startup that it was built
for the card it runs on. Validated on doplxyz's machine with both cards. The live server tests pass on each.
Two AMD cards also run as one layer split from a hand-written config (setup installs one AMD card). It gives
exactly the output of one card when the experts sit in the same place; see docs/AMD_HIP.md.
The server can give the GPU back (#208, @bytethecookie).
--idle-unload 600unloads the model after 600 s without requests; the next request loads it again.POST /unloadandPOST /loaddo it on demand.--min-free-vram-mib Nloads only when that much VRAM is free (else 503), and--before-load "cmd"runs a
command first (e.g. one that unloads another server's model).
Low-RAM mode, resident variant. When the experts don't fit the RAM but the ones the GPU doesn't hold do,
setup now copies exactly those into RAM at start (--resident-experts) instead of reading them from the file.
The same output. On a cold file cache the prompt went from 150 to 1,197 tok/s and answers were 26% faster.
Opt-in and experimental (off unless you turn them on):
- Contexts past the trained 262K (#84, @j-luwierski):
--rope-scaling yarn|linearand--rope-scale F,
or setup's--contextpast 262144. Experimental: at 320K, facts planted at 300K were found at 10%, 50% and
90% depth. Seedocs/DETAILS.mdfor the measured quality. - Several conversations kept at once (#189, @jeremiahritchey):
--conversation-cache-mib 8192 --conversation-cache-slots 4parks up to four conversations in RAM, so switching back doesn't re-read the
prompt. Not yet with a layer split. With the option off, the engine is byte-identical to 0.1.29. - Greedy output independent of drafting (#152, reported by @clapbr):
STRATA_IQ_MT_MIN=1in the config's
envmakes the IQ models' CPU experts round the same whatever the verify window held. Costs 1-3% decode on
IQ3_S (AVX-512); the default is unchanged. - Coupled draft sampling:
STRATA_SPEC_COUPLED=1samples drafts and the verifier from one random stream.
No clear gain in our tests, so it stays off. - Docker (#96, @djmaze):
docker build -t strata .builds the engine in the image. The server is PID 1 and
stops cleanly on SIGTERM; KV, GPUs and low-RAM mode are set with environment variables. See the README. - Pascal cards (#124, @ruibeikaa): the engine builds and runs on compute capability 6.0 (P100) with exact
fallbacks for__dp4aand__nanosleep. Not in the ready-made engine; compile it yourself. Not tested here
(no Pascal card). - Shared expert arena (#129, @rhgo1749): Linux:
--shared-expert-arena FILEin/dev/shmlets several
engines on one machine share the RAM copy of the experts.
Checked before the release:
-
Byte-identical to 0.1.29 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),
including the prompt path's internal state, with the same decode speed. -
Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:
- a fact found in a 57,566-token document;
- the follow-up turn reusing the whole prompt;
- a tool call;
- a cancelled long prompt followed by a new request;
- a sampled answer.
None of them restarted the engine or ran low on VRAM. Prompt speed and free VRAM were the same as 0.1.29.
-
The conversation cache (#189): with it off, the same output as 0.1.29 up to 128K; with it on, the reuse,
memory-pressure and admission tests pass. -
Rope scaling (#84): the parity tests, and needles at 300K in a 320K context.
-
The server's tests (80) and the CPU expert parity test (0 failures; with
STRATA_IQ_MT_MIN=1no row differs between
one token and a group). -
Linux (WSL, RTX 5070): Q2_0 byte-identical to 0.1.29 with a fixed cache (10/10), and the default settings pass.
-
AMD (RX 9070 XT and R9700): the HIP build with its tests (45 of 47; the 2 others need a model fixture or an
AVX-512 CPU), the live server test on each card, and the two-card layer split.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.30.
The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.