Repository navigation
Strata v0.1.17
Claude Code works with Strata, sampling matches llama.cpp, and you can choose the GPU.
- Claude Code can use Strata (#55, #56, thanks @Nicolas0315): the server accepts
/v1/messages?beta=true, and a system message in the middle of a conversation (Claude Code sends its hook context that way; some OpenAI clients send late "developer" messages) no longer breaks the chat template. SetANTHROPIC_BASE_URL=http://127.0.0.1:8080andANTHROPIC_MODELto a Claude model name it knows (Strata ignores the name) - see the details. - Fixed: sampling applied the repetition, frequency and presence penalties twice (#53, thanks @hendrikp). With
presence_penalty1.5 at temperature 0.7, a repeated word was about 4x less likely than it should be. The order is now llama.cpp's: penalties once, then top_k, top_p, min_p, temperature. Greedy answers (temperature 0, the default) are unchanged. - Choose the GPU on a PC with more than one (#51, thanks @ivancheg8): setup uses the one with the most VRAM, or
START-HERE.bat --gpu 1(numbered asnvidia-smishows them); the Monitor shows that card. tools/needle_bench.pychecks long-context recall through a running server: it hides a code word in a long text and asks for it (#33).
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat with the model window closed: it updates the engine to 0.1.17 by itself. On Linux, run ./setup.sh: it compiles the new engine.
The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.