Strata v0.1.20
Faster new chats for agent clients, a hit-rate column in the Monitor, a PCIe-aware default, and community fixes.
- New chats reuse the system prompt (#62 by @j-luwierski, #65 by @code-martin):
- A prompt read from the start is now also checkpointed at the end of its system prompt (with the tool list), when that is 2,048 tokens or more.
- That checkpoint is kept while the others rotate by least recent use, so the next new chat of the same client reads only what comes after it.
- Measured: a 3,883-token system prompt is reused whole by a new chat, also after more turns than the cache has slots.
--prompt-cache-root 0turns it off. See the details.
- Monitor: expert-cache hit rate per request (#69 by @code-martin): the share of the experts looked up while writing the answer that were already in VRAM.
- A default PCIe share that fits your link (#44 by @pipeob0):
- The engine measures the PCIe bandwidth at start. On a narrower link (e.g. an x8 card) it lowers the share of the missing experts that the GPU reads over PCIe.
- x16 links keep the measured default.
--calibratestill measures it outright.
- OrcaRouter's Uncensored IQ3_XXS can be packed with
iq_pack.py --compat-bf16(#67 by @acrogenesis); see docs/ORCA.md. It is a manual setup, not an installer option. The official models pack exactly as before. - Fixed:
- The parity tool fails the run on a kernel mismatch instead of only printing it (#64 by @pipeob0).
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.20 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.