Skip to content

Strata v0.1.36

Choose a tag to compare

@Niko1221 Niko1221 released this 02 Oct 16:38
· 285 commits to main since this release

Faster: Q2_0 reads prompts 16-22% faster and writes up to 18% faster at long contexts. Plus one-click updates (UPDATE.bat), an opt-in memory of which experts you use, and clearer logs.

Speed (#136), measured on an RTX 5070 12 GB, 0.1.35 -> 0.1.36:

Q2_0 Reads your prompt Writes answers
4K prompt 1,294 -> 1,570 tokens/s (+21%) 89 -> 93.5 (+5%)
32K prompt 2,170 -> 2,653 (+22%) 80 -> 81
128K prompt 2,123 -> 2,468 (+16%) 64.5 -> 76.4 (+18%)
  • New prompt kernels for Q2_0 (RTX 30 and newer): the experts are multiplied by fused int8 tensor-core kernels
    that read each expert where it already is (no copies), with the grouping done on the GPU. They are as accurate as
    before: measured against an FP16 reference over 2,000 teacher-forced tokens after a long document, as close as the
    previous kernels at 8K and closer at 32K (KL 0.009 vs 0.012, perplexity 4.121 vs 4.169, the reference 4.129).
    Answers are not byte-identical to 0.1.35 any more, the same kind of change as a new GPU kernel's rounding;
    STRATA_PF_FUSED=0 in the config's env brings the previous kernels back.
  • Faster long-context decoding on RTX 50 cards: the attention's block selection and the greedy token pick run on
    thread-block clusters. The output is exactly the same (bit-identical tested). Older cards keep the previous
    kernels.
  • Opt-in for the IQ sizes: STRATA_PF_FUSED=1 also runs fused kernels for IQ2_XS / IQ3_XXS / IQ3_S (IQ2_XS
    prompts +12% at 4K, +3% at 32K; the IQ3 sizes about even, so they stay off by default).
  • The approach was inspired by the measurements of another engine for this model,
    flashrt; Strata's kernels are its own.

UPDATE.bat / update.sh (#475): updates Strata without starting the model. In a git clone it pulls the new
files (a zip copy is told to download the new zip), then updates the Python packages, the engine, each model's
settings and the draft vocabulary. Then it stops.

Keep what the expert cache learned (#477, opt-in): with "expert_profile_save": "path\\to\\my-profile.bin" in
strata-<model>.json (or --expert-profile-save), the engine saves which experts your work uses (every 10 minutes
and on exit), and the next start begins from it instead of the shipped profile. Off by default.

Fixes:

  • A prompt cancelled while it is being read (#471): the log and /metrics say how much was read, with the real
    reading speed.
  • "the draft head does not fit" (#474): the engine says what it needs, the free VRAM, and the smaller
    --draft-vocab cyrillic / en options with their sizes; the server shows the hint.

Checked before the release:

  • The same answers as 0.1.35 on IQ3_XXS, IQ3_S and the Coder (10/10 each), and on Q2_0 with STRATA_PF_FUSED=0
    (10/10). Q2_0's default answers changed by design (the quality check above).
  • The new kernels' tests: the fused kernels against an FP32 reference and the previous kernels for every
    format (the same error), the cluster kernels bit-identical to the previous ones over 650 cases up to 262K, the
    existing parity tests.
  • Speed: the 5-pair A/B on Q2_0 and IQ3_S (+1.9 / +0.6 / +0.6 / +0.7%, the same expert slots) and the table above.
  • Real use at a 57K-token prompt through the server on Q2_0 and the Coder; the fused path also through serve
    sessions with cancels, KV streaming and mixed chunk sizes.
  • Linux (WSL, RTX 5070): Q2_0 identical to 0.1.35 with STRATA_PF_FUSED=0 (10/10), and the defaults pass.
  • AMD: the Windows HIP zip builds (AMD keeps its own kernels; the new ones are NVIDIA-only).
  • Tests: the tools' and setup's (261), the server's (158).

Updating: run UPDATE.bat from now on. This time: git pull (or download the new zip), then START-HERE.bat
(Linux: ./setup.sh). Setup installs engine 0.1.36.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, the HIP runtime
    next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).