Skip to content

Strata v0.1.30

Choose a tag to compare

@Niko1221 Niko1221 released this 30 Sep 17:50
· 502 commits to main since this release

Short prompts up to 28% faster (the same answers), multi-GPU and AMD RDNA4 improvements, a server that can give the GPU back, and several opt-in features.

Faster short prompts. Prompts of 1K to 4K tokens read all experts from the stream from 1,024-token chunks
(0.1.29: 2,048): +17-28% prompt speed on this RTX 5070, with the same output (suggested in the issues).
STRATA_PREFILL_STREAM_MIN overrides the threshold.

Multi-GPU (#216, @gopinath87607). In a layer split:

  • Each card now keeps only its own layers' session state, which saved 0.93 GiB on a later card in the author's
    4-GPU test.
  • Each card lends the prompt path the tail of its own expert cache, instead of every card after the first
    holding back 1 GiB for the whole session. --prefill auto can then pick larger chunks (8,192 tokens instead of
    2,048 there).
  • The Monitor's expert cache counts every card (it showed only the first).
  • Snapshots and checkpoints follow the split: a card saves and restores only its own layers.

Single-GPU runs are unchanged (byte-identical).

AMD RDNA4 (#178, @doplxyz). The RX 9070 / 9070 XT and the Radeon AI PRO R9700 (gfx1201) are supported beside
the RX 7900 series: setup picks ROCm and the settings per card, and the engine checks at startup that it was built
for the card it runs on. Validated on doplxyz's machine with both cards. The live server tests pass on each.
Two AMD cards also run as one layer split from a hand-written config (setup installs one AMD card). It gives
exactly the output of one card when the experts sit in the same place; see docs/AMD_HIP.md.

The server can give the GPU back (#208, @bytethecookie).

  • --idle-unload 600 unloads the model after 600 s without requests; the next request loads it again.
  • POST /unload and POST /load do it on demand.
  • --min-free-vram-mib N loads only when that much VRAM is free (else 503), and --before-load "cmd" runs a
    command first (e.g. one that unloads another server's model).

Low-RAM mode, resident variant. When the experts don't fit the RAM but the ones the GPU doesn't hold do,
setup now copies exactly those into RAM at start (--resident-experts) instead of reading them from the file.
The same output. On a cold file cache the prompt went from 150 to 1,197 tok/s and answers were 26% faster.

Opt-in and experimental (off unless you turn them on):

  • Contexts past the trained 262K (#84, @j-luwierski): --rope-scaling yarn|linear and --rope-scale F,
    or setup's --context past 262144. Experimental: at 320K, facts planted at 300K were found at 10%, 50% and
    90% depth. See docs/DETAILS.md for the measured quality.
  • Several conversations kept at once (#189, @jeremiahritchey): --conversation-cache-mib 8192 --conversation-cache-slots 4 parks up to four conversations in RAM, so switching back doesn't re-read the
    prompt. Not yet with a layer split. With the option off, the engine is byte-identical to 0.1.29.
  • Greedy output independent of drafting (#152, reported by @clapbr): STRATA_IQ_MT_MIN=1 in the config's
    env makes the IQ models' CPU experts round the same whatever the verify window held. Costs 1-3% decode on
    IQ3_S (AVX-512); the default is unchanged.
  • Coupled draft sampling: STRATA_SPEC_COUPLED=1 samples drafts and the verifier from one random stream.
    No clear gain in our tests, so it stays off.
  • Docker (#96, @djmaze): docker build -t strata . builds the engine in the image. The server is PID 1 and
    stops cleanly on SIGTERM; KV, GPUs and low-RAM mode are set with environment variables. See the README.
  • Pascal cards (#124, @ruibeikaa): the engine builds and runs on compute capability 6.0 (P100) with exact
    fallbacks for __dp4a and __nanosleep. Not in the ready-made engine; compile it yourself. Not tested here
    (no Pascal card).
  • Shared expert arena (#129, @rhgo1749): Linux: --shared-expert-arena FILE in /dev/shm lets several
    engines on one machine share the RAM copy of the experts.

Checked before the release:

  • Byte-identical to 0.1.29 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),
    including the prompt path's internal state, with the same decode speed.

  • Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:

    • a fact found in a 57,566-token document;
    • the follow-up turn reusing the whole prompt;
    • a tool call;
    • a cancelled long prompt followed by a new request;
    • a sampled answer.

    None of them restarted the engine or ran low on VRAM. Prompt speed and free VRAM were the same as 0.1.29.

  • The conversation cache (#189): with it off, the same output as 0.1.29 up to 128K; with it on, the reuse,
    memory-pressure and admission tests pass.

  • Rope scaling (#84): the parity tests, and needles at 300K in a 320K context.

  • The server's tests (80) and the CPU expert parity test (0 failures; with STRATA_IQ_MT_MIN=1 no row differs between
    one token and a group).

  • Linux (WSL, RTX 5070): Q2_0 byte-identical to 0.1.29 with a fixed cache (10/10), and the default settings pass.

  • AMD (RX 9070 XT and R9700): the HIP build with its tests (45 of 47; the 2 others need a model fixture or an
    AVX-512 CPU), the live server test on each card, and the two-card layer split.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.30.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.