Strata v0.1.13
Strata engine 0.1.13 - long prompts about twice as fast
Update: run START-HERE.bat (Windows) or git pull + ./setup.sh (Linux). Setup installs the new engine and
switches your config to the new prompt reading (--prefill auto) by itself.
Faster prompts
Reading a prompt (the time before the first word) is about 2x faster on long prompts. RTX 5070 12 GB, 64 GB RAM:
| 0.1.12 | 0.1.13 | |
|---|---|---|
| Q2_0, 32K-token prompt | 572 tokens/s | 1,290 tokens/s |
| IQ3_S, 32K-token prompt | 383 tokens/s | 1,208 tokens/s |
| Through the server, 7K-token prompt (Q2_0, 128K context) | 529 tokens/s | 1,077 tokens/s |
| Through the server, 28K-token prompt | 584 tokens/s | 1,249 tokens/s |
How: bigger prompt chunks (up to 8,192 tokens, sized to what fits), llama.cpp's quantized MMQ kernels for the
experts, the next layer's experts streaming over PCIe while the current layer runs, and the PLE block computed for the
whole chunk at once. Output speed is unchanged; needle tests pass 5/5 from 1K to 262K. Details:
bench/results/2026-09-28-prefill-speed.
Stalls (issues #29, #31)
- A request that stops moving now ends after 60 s (was 120 s), and the log gets a stall report under the
no progressline: what every expert-pool thread and the GPU handshake were doing, memory and page faults. On
Windows the engine also writesstrata-stall-<pid>.dmpwith every thread's stack. If you see one, please attach
both to the issue. - The engine prints its version at start (
session is up (engine 0.1.13)). - Linux:
git pull+./setup.shnow recompiles a self-compiled engine when its source changed (it kept the old one).
Assets
strata-windows-x64.zip- the engine (RTX 30/40/50).strata-windows-x64-symbols.zip- the matchingstrata.pdb, only needed to read a stall dump.