Skip to content

Strata v0.1.17

Choose a tag to compare

@Niko1221 Niko1221 released this 28 Sep 13:40
· 1064 commits to main since this release

Claude Code works with Strata, sampling matches llama.cpp, and you can choose the GPU.

  • Claude Code can use Strata (#55, #56, thanks @Nicolas0315): the server accepts /v1/messages?beta=true, and a system message in the middle of a conversation (Claude Code sends its hook context that way; some OpenAI clients send late "developer" messages) no longer breaks the chat template. Set ANTHROPIC_BASE_URL=http://127.0.0.1:8080 and ANTHROPIC_MODEL to a Claude model name it knows (Strata ignores the name) - see the details.
  • Fixed: sampling applied the repetition, frequency and presence penalties twice (#53, thanks @hendrikp). With presence_penalty 1.5 at temperature 0.7, a repeated word was about 4x less likely than it should be. The order is now llama.cpp's: penalties once, then top_k, top_p, min_p, temperature. Greedy answers (temperature 0, the default) are unchanged.
  • Choose the GPU on a PC with more than one (#51, thanks @ivancheg8): setup uses the one with the most VRAM, or START-HERE.bat --gpu 1 (numbered as nvidia-smi shows them); the Monitor shows that card.
  • tools/needle_bench.py checks long-context recall through a running server: it hides a code word in a long text and asks for it (#33).

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.17 by itself. On Linux, run ./setup.sh: it compiles the new engine.

The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.