Skip to content

Strata v0.1.22

Choose a tag to compare

@Niko1221 Niko1221 released this 29 Sep 09:41
· 62 commits to main since this release

Faster prompts, and several GPUs without typing any flags.

Prompts are read faster (RTX 5070, Q2_0): a 32K-token prompt 1,213 → 1,646 tokens/s, a 128K prompt with KV streaming 1,144 → 1,596 tokens/s (time to first token 116 s → 83 s). Where it came from:

  • the sparse attention over a prompt runs on tensor cores (4x faster at FP32-level accuracy; needle tests pass at every depth);
  • the n-gram table rows are read from disk by several threads, on Windows and Linux (on Linux the reads were one at a time: a 16K prompt read 3x faster in our test);
  • the selection, indexer, short-conv and GDN kernels of the prompt path, the embedding gather, the draft layer's prompt pass and the expert streaming of the IQ quants no longer wait on the host (these are bit-identical: with STRATA_PROMPT_ATTN_OLD=1 every output is byte-identical to 0.1.21);
  • the expert-cache slots the prompt path borrows are given back in one batch.
    Output speed is unchanged.

Two or three GPUs (#112, #77, #95, #36; thanks @megakilo, @sal7g86, @pdsmike):

  • START-HERE.bat / ./setup.sh lists your NVIDIA cards, says which ones Strata can use (and why not, e.g. older than the RTX 30 series), and asks which to use, with the two best together recommended.
  • A model you installed on one card asks once, at its next start, whether to use both from now on.
  • --gpus 0,1, --gpus all or --gpu 0,1 work at any start and are remembered; --gpu N runs one card for that start.
  • Fixed: 0.1.21's --setup --gpus 0,1 stopped at the end with "there is no GPU [0, 1]"; --gpus on an installed model started on one card; an update unzipped next to a split install stopped at step 1.
  • Images, the experimental speed projection (control vectors), KV streaming and mid-prompt checkpoints now work across cards; --calibrate measures the split; the Monitor shows every card (total VRAM, power and load, each card's own reading underneath).
  • The engine is compiled for every chosen card's generation when you build it yourself.
  • Tested on an RTX 5080 + RTX 3090 (+ a 2080 Ti, which is refused with its reason). Linux uses the same setup code, but we could not test several GPUs on Linux; reports welcome.

Also:

  • Setup compiles with Visual Studio 2022 even when Visual Studio 2026 is installed too (CUDA 13 refuses 2026).
  • The streaming detokenizer no longer slows down as a reply grows (it cost 4 ms per token after 16K tokens).

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.22.

The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.