Skip to content

Strata v0.1.19

Choose a tag to compare

@Niko1221 Niko1221 released this 28 Sep 17:13
· 168 commits to main since this release

Penalties that really apply, tuning for your PC, tools from MCP servers, and a fix for starting with a small page file.

  • Fixed: repetition, presence and frequency penalties were mostly not applied (#53).
    • What went wrong: speculative decoding checks several guessed tokens at once, but only the first of each batch got the request's penalty history. The others were judged against memory that was never filled in.
    • What it meant: with ~3 tokens per step, most tokens ignored the penalties, and penalty_last_n near 4096 could crash the engine.
    • Now every checked token gets its own history, exactly as if the engine decoded one token at a time.
    • The cost: requests with penalties are 1-11% slower, because more draft guesses are rejected. Requests without penalties give the same answers at the same speed.
    • Found by an engineering review of the engine. Tested on Q2_0, IQ2_XS (Swift), IQ3_XXS, IQ3_S and the Coder.
  • Sampler:
    • top_k = 0 or above 64 now means "the widest list the engine keeps" (64), instead of an unwritten token (engine) or a silent 20 (server);
    • token ids outside the vocabulary in a history are ignored.
  • strata.exe on the command line samples again: with --temperature/--seed the speculative loop used to decode greedily. It also prints the sampling it uses.
  • The model starts with a small Windows page file (#60, thanks @shanoor).
    • Why it failed: memory on the graphics card also needs room in Windows' commit (RAM plus page file). With the page file off or tiny, the expert cache could not be allocated even though VRAM was free, and the engine stopped.
    • Now it retries with a smaller cache and says why.
    • Setup warns when the page file is under 4 GB. The fix: set it to "System managed".
  • START-HERE.bat --calibrate: tuning for your PC.
    • It measures three engine settings that depend on the machine: the PCIe share, the draft depth and the CPU threads.
    • It keeps a setting only when it is more than 3% faster, and remembers the result per PC and model.
    • Setup offers it after an install. On the RTX 5070 test PC it made the Coder 7.6% faster.
  • Tools from MCP servers in the chat page.
    • List servers in the run config under mcp_servers; Claude Desktop's mcpServers block can be pasted as it is.
    • The model can then call their tools in the web app's chat: files, search, GitHub and so on. Each call shows as a card.
    • API clients are not affected. See the details.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.19 by itself. On Linux, run ./setup.sh: it compiles the new engine.

The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.