Strata v0.1.21
One model across two or three NVIDIA cards (experimental).
- Layer split across GPUs:
- Setup:
START-HERE.bat --setup --gpus 0,2(Linux:./setup.sh --setup --gpus 0,2), or"gpu": [0, 2]in a run config. - Each card runs a range of the model's layers and keeps the experts of those layers in its own VRAM; the last card also runs the output head and the draft layer.
- Prompts flow through the cards in a pipeline: the next card reads chunk c while the first reads chunk c+1.
- Conversation checkpoints, adaptive expert swaps and the PCIe share work per card.
- No NVLink or peer-to-peer access needed: cards in x4 slots work.
- Setup:
- Measured on an RTX 5080 + RTX 3090 (Ryzen 9 9950X3D) with the Coder, 32K context:
- prompts 18-20% faster than the 5080 alone: 2,357 vs 1,970 tokens/s at 28K tokens;
- decoding on par (84 / 110 tokens/s, story / code), because both caches then hold ~99% of the routed experts and the per-layer GPU time decides.
- Placement:
--layer-split auto(the default with several cards) places the split from each card's free VRAM and speed. It picked a split within 0-12% of the best one measured.- List the fastest card first, and leave out a much slower one: an RTX 2080 Ti as a third card made the pair slower.
- Details, limits and all measurements: docs/MULTI_GPU.md and
bench/results/2026-09-29-layer-split/.
- Also included:
- The experimental helper-GPU expert caches from #16 by @Daerdaal (
--expert-cache-remote, docs/SECOND_GPU.md): the first multi-GPU groundwork. - On a PCIe link below 4 GB/s (x1 risers), no missing expert is read over PCIe any more; the CPU computes them all (it is faster there).
-DSTRATA_EXPERIMENTAL_SM75=ONbuilds the engine for an RTX 20-series card (Turing), meant as an extra layer-split card. Not supported otherwise.
- The experimental helper-GPU expert caches from #16 by @Daerdaal (
- One GPU is unchanged:
- Tested against 0.1.20 on Windows with every model size (Coder, Q2_0, IQ3_XXS, IQ3_S, Swift IQ2_XS) and on Linux (Q2_0): the same answers, byte for byte.
- Same speed.
Thanks to the two community members who lent their PCs for the multi-GPU testing.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.21 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.