Skip to content

Strata v0.1.13

Choose a tag to compare

@Niko1221 Niko1221 released this 28 Sep 05:09
· 82 commits to main since this release

Strata engine 0.1.13 - long prompts about twice as fast

Update: run START-HERE.bat (Windows) or git pull + ./setup.sh (Linux). Setup installs the new engine and
switches your config to the new prompt reading (--prefill auto) by itself.

Faster prompts

Reading a prompt (the time before the first word) is about 2x faster on long prompts. RTX 5070 12 GB, 64 GB RAM:

0.1.12 0.1.13
Q2_0, 32K-token prompt 572 tokens/s 1,290 tokens/s
IQ3_S, 32K-token prompt 383 tokens/s 1,208 tokens/s
Through the server, 7K-token prompt (Q2_0, 128K context) 529 tokens/s 1,077 tokens/s
Through the server, 28K-token prompt 584 tokens/s 1,249 tokens/s

How: bigger prompt chunks (up to 8,192 tokens, sized to what fits), llama.cpp's quantized MMQ kernels for the
experts, the next layer's experts streaming over PCIe while the current layer runs, and the PLE block computed for the
whole chunk at once. Output speed is unchanged; needle tests pass 5/5 from 1K to 262K. Details:
bench/results/2026-09-28-prefill-speed.

Stalls (issues #29, #31)

  • A request that stops moving now ends after 60 s (was 120 s), and the log gets a stall report under the
    no progress line: what every expert-pool thread and the GPU handshake were doing, memory and page faults. On
    Windows the engine also writes strata-stall-<pid>.dmp with every thread's stack. If you see one, please attach
    both to the issue.
  • The engine prints its version at start (session is up (engine 0.1.13)).
  • Linux: git pull + ./setup.sh now recompiles a self-compiled engine when its source changed (it kept the old one).

Assets

  • strata-windows-x64.zip - the engine (RTX 30/40/50).
  • strata-windows-x64-symbols.zip - the matching strata.pdb, only needed to read a stall dump.