Skip to content

Strata v0.1.29

Choose a tag to compare

@Niko1221 Niko1221 released this 30 Sep 12:18
· 356 commits to main since this release

Sampled answers up to 40% faster (the same text), faster prompt kernels, correctness fixes, and several contributions.

Sampled answers (#197, @gputier). With a temperature above 0 (most chat apps), each token's top-k selection now
runs split across the whole GPU, and the rest of the sampling runs on one warp. It picks exactly the same tokens:
with the same seed, the output is identical to 0.1.28 (checked on 12 runs across 3 sampling settings, and by
sampler_parity against the old kernel). On an RTX 5070:

  • Q2_0: +4% with top-k 20, +15-20% with top-k 40, +38-42% with top-k 64.
  • The Coder: +4% to +26% on the same settings.
  • Greedy answers (temperature 0) are unchanged.

Faster kernels, the same bits:

  • The QSA block scores read each key block once for all of a verify window's queries (#187, @q8atnight).
  • The GDN recurrence of the prompt path loads the next token's inputs while the current one computes (#188,
    @q8atnight).
  • AVX2 CPUs prefetch the expert rows ahead, as the AVX-512 kernels already did (#207, @pipeob0).

Correctness (#154, @gputier):

  • A NaN stays a NaN when converted to bf16, as in ggml. It used to become -0 or infinity and hide the error.
  • The q8 split GEMV kernel reaches its barrier before a warp without rows returns.
  • A KV-streaming block that could not be made resident is masked instead of read from before its pool.
  • The verify window's tables are guarded at their size.

Also:

  • Windows: the engine, the image encoder and MCP servers now end with the server, however it ends (closing the
    window, Task Manager) (#181, @Apposite245).
  • The dashboard shows the prompt speed beside the answer speed (#163, @hendrikp).
  • --ple-io ram keeps the PLE table locked in RAM on Linux (#202, @q8atnight).
  • WSL2 with 3 GPUs: an opt-in for the driver's ~1 GiB pinned budget, and the token embedding falls back to VRAM when
    pinning fails (#158, @MrRoza).
  • Setup: the 3-bit models' 262K context is counted against the RAM instead of a fixed 90 GB, so a PC with ~84 GB
    keeps 262K (#134, @architectds).
  • STRATA_PREFILL_RING=8 selects routed-only staging for large chunks (#225, @rluisr).
  • CLI generate zeroes the session state before the prompt for every pack (#167, @samuelishida).
  • A CUDA fault in the prompt path now ends the engine at once instead of after the 60 s watchdog (#224).
  • The build warns when RTX 50 code is compiled with CUDA older than 13.0 (#220, #224).
  • The pack tool leaves no partial pack behind when it stops (#172).
  • A bitwise test of the verify window's multi-token GEMV contract (#148, @enkynakamura).

Checked before the release:

  • Byte-identical to 0.1.28 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),
    including the prompt path's internal state.

  • Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:

    • a fact found in a 58,700-token document;
    • the follow-up turn reusing the whole prompt;
    • a tool call;
    • a cancelled long prompt followed by a new request;
    • a sampled answer.

    None of them restarted the engine or ran low on VRAM.

  • The parity tests (sampler in all three modes, bf16, element-wise, s_gemv, the new multi-token test) and the
    server's tests (71).

  • Builds and runs the same on Linux.

  • AMD: not tested again for this release (our AMD test PC was not reachable).

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.29.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.