v0.1.0-alpha
Pre-release
Pre-release
First public alpha of runner — a compact local LLM inference engine in plain C (GGUF, CPU/CUDA/Metal, OpenAI-compatible server, sampler-level JSON-schema enforcement).
Binaries: Linux x86_64 and Windows x86_64 (both need AVX2), macOS arm64 (Apple Silicon). Or build from source: make — no dependencies beyond a C compiler.
What to test: run your GGUF models on your hardware. If anything crashes, misbehaves, or underperforms, open an issue and include the output of runner --version and runner --caps.