v0-pre
Pre-release
Pre-release
Kolosal CLI v0.1 is here to make you run and deploy any local LLM easily.
- Single Binary: One lightweight executable, install anywhere instantly
- Universal GPU Support: Powered by llama.cpp + Vulkan—works on every GPU
- Auto-Scaling: Models scale down when idle, seamless switching
- Instant API: Every model available at localhost:8080 (OpenAI compatible)
- Smart Memory: Built-in approximator shows what models fit your hardware
- Hugging Face Ready: Run any GGUF model with one command