Strata v0.1.24
Long prompts read faster: the sparse attention's selection runs on tensor cores.
For every token of a prompt, Strata picks which earlier parts of the context its attention reads (the QSA selection). At long context that choice was 18% of the prompt time. Now:
- the scores it chooses by are computed on tensor cores, at FP32-level accuracy (3xTF32);
- the top-k choice keeps each query's scores in registers - it picks exactly the same parts as before (also while writing answers).
RTX 5070, Q2_0: a 128K-token prompt 1,608 → 1,843 tokens/s (time to first token 82.7 s → 72.3 s), a 32K prompt 1,742 → 1,797 tokens/s. The selection step itself is 3.3-3.9x faster.
Checked before the release: needle tests 15/15 (8K, 16K, 32K at five depths); with the old selection switched on (STRATA_SELECT_OLD=1 STRATA_TOPK_OLD=1) the output is byte-identical to 0.1.23 on Q2_0 and the Coder, including the prompt path's internal state at 4K and 20K; a live server test of every endpoint and image requests passes; builds and runs the same on Linux. Answers can differ from 0.1.23 in the last bits of a score (another summation order), the same kind of change as 0.1.22's attention kernel.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.24.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.