Strata v0.1.15
Fixed: IQ3_XXS stopping on long prompts, and new speed tables.
- Fixed: IQ3_XXS could stop on prompts of 64K tokens and more on a 12 GB card, with
out of memory: cudaFuncSetAttributein the log. CUDA loads a kernel's code the first time it is used, and in the middle of a long prompt there was no VRAM left for it. The engine now loads all of its kernels when it starts, before the expert cache takes the free VRAM (~30 MB, about 20 fewer experts on the GPU). Other models and shorter prompts were not affected. - New speed tables in the README and details, measured with the current engine and settings: prompts are read 1.2x (1K) to 2.4x (32K-128K) faster than the old tables showed - Q2_0 reads a 32K prompt at 1,308 tokens/s (was 571), IQ3_S at 1,070 (was 397). Raw numbers:
bench/results/2026-09-28-speed-0114.
Updating: get the latest files (git pull or download the repo again), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.15 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.