Skip to content

Strata v0.1.21

Choose a tag to compare

@Niko1221 Niko1221 released this 29 Sep 04:33
· 101 commits to main since this release

One model across two or three NVIDIA cards (experimental).

  • Layer split across GPUs:
    • Setup: START-HERE.bat --setup --gpus 0,2 (Linux: ./setup.sh --setup --gpus 0,2), or "gpu": [0, 2] in a run config.
    • Each card runs a range of the model's layers and keeps the experts of those layers in its own VRAM; the last card also runs the output head and the draft layer.
    • Prompts flow through the cards in a pipeline: the next card reads chunk c while the first reads chunk c+1.
    • Conversation checkpoints, adaptive expert swaps and the PCIe share work per card.
    • No NVLink or peer-to-peer access needed: cards in x4 slots work.
  • Measured on an RTX 5080 + RTX 3090 (Ryzen 9 9950X3D) with the Coder, 32K context:
    • prompts 18-20% faster than the 5080 alone: 2,357 vs 1,970 tokens/s at 28K tokens;
    • decoding on par (84 / 110 tokens/s, story / code), because both caches then hold ~99% of the routed experts and the per-layer GPU time decides.
  • Placement:
    • --layer-split auto (the default with several cards) places the split from each card's free VRAM and speed. It picked a split within 0-12% of the best one measured.
    • List the fastest card first, and leave out a much slower one: an RTX 2080 Ti as a third card made the pair slower.
    • Details, limits and all measurements: docs/MULTI_GPU.md and bench/results/2026-09-29-layer-split/.
  • Also included:
    • The experimental helper-GPU expert caches from #16 by @Daerdaal (--expert-cache-remote, docs/SECOND_GPU.md): the first multi-GPU groundwork.
    • On a PCIe link below 4 GB/s (x1 risers), no missing expert is read over PCIe any more; the CPU computes them all (it is faster there).
    • -DSTRATA_EXPERIMENTAL_SM75=ON builds the engine for an RTX 20-series card (Turing), meant as an extra layer-split card. Not supported otherwise.
  • One GPU is unchanged:
    • Tested against 0.1.20 on Windows with every model size (Coder, Q2_0, IQ3_XXS, IQ3_S, Swift IQ2_XS) and on Linux (Q2_0): the same answers, byte for byte.
    • Same speed.

Thanks to the two community members who lent their PCs for the multi-GPU testing.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.21 by itself. On Linux, run ./setup.sh: it compiles the new engine.

The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.