Strata v0.1.29
Sampled answers up to 40% faster (the same text), faster prompt kernels, correctness fixes, and several contributions.
Sampled answers (#197, @gputier). With a temperature above 0 (most chat apps), each token's top-k selection now
runs split across the whole GPU, and the rest of the sampling runs on one warp. It picks exactly the same tokens:
with the same seed, the output is identical to 0.1.28 (checked on 12 runs across 3 sampling settings, and by
sampler_parity against the old kernel). On an RTX 5070:
- Q2_0: +4% with top-k 20, +15-20% with top-k 40, +38-42% with top-k 64.
- The Coder: +4% to +26% on the same settings.
- Greedy answers (temperature 0) are unchanged.
Faster kernels, the same bits:
- The QSA block scores read each key block once for all of a verify window's queries (#187, @q8atnight).
- The GDN recurrence of the prompt path loads the next token's inputs while the current one computes (#188,
@q8atnight). - AVX2 CPUs prefetch the expert rows ahead, as the AVX-512 kernels already did (#207, @pipeob0).
- A NaN stays a NaN when converted to bf16, as in ggml. It used to become -0 or infinity and hide the error.
- The q8 split GEMV kernel reaches its barrier before a warp without rows returns.
- A KV-streaming block that could not be made resident is masked instead of read from before its pool.
- The verify window's tables are guarded at their size.
Also:
- Windows: the engine, the image encoder and MCP servers now end with the server, however it ends (closing the
window, Task Manager) (#181, @Apposite245). - The dashboard shows the prompt speed beside the answer speed (#163, @hendrikp).
--ple-io ramkeeps the PLE table locked in RAM on Linux (#202, @q8atnight).- WSL2 with 3 GPUs: an opt-in for the driver's ~1 GiB pinned budget, and the token embedding falls back to VRAM when
pinning fails (#158, @MrRoza). - Setup: the 3-bit models' 262K context is counted against the RAM instead of a fixed 90 GB, so a PC with ~84 GB
keeps 262K (#134, @architectds). STRATA_PREFILL_RING=8selects routed-only staging for large chunks (#225, @rluisr).- CLI
generatezeroes the session state before the prompt for every pack (#167, @samuelishida). - A CUDA fault in the prompt path now ends the engine at once instead of after the 60 s watchdog (#224).
- The build warns when RTX 50 code is compiled with CUDA older than 13.0 (#220, #224).
- The pack tool leaves no partial pack behind when it stops (#172).
- A bitwise test of the verify window's multi-token GEMV contract (#148, @enkynakamura).
Checked before the release:
-
Byte-identical to 0.1.28 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),
including the prompt path's internal state. -
Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:
- a fact found in a 58,700-token document;
- the follow-up turn reusing the whole prompt;
- a tool call;
- a cancelled long prompt followed by a new request;
- a sampled answer.
None of them restarted the engine or ran low on VRAM.
-
The parity tests (sampler in all three modes, bf16, element-wise, s_gemv, the new multi-token test) and the
server's tests (71). -
Builds and runs the same on Linux.
-
AMD: not tested again for this release (our AMD test PC was not reachable).
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.29.
The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.