Janus is a single Go binary that runs .gguf models on your machine (GPU or CPU) and exposes an OpenAI-compatible API. No Python, no Docker, no Ollama required.
Use it your way: call it from the command line (curl, PowerShell, scripts), wire it into Cursor / Cline / any OpenAI client — same local models, whatever workflow fits you.
- Local inference — llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
- OpenAI-compatible API —
/v1/chat/completions,/v1/models - Hot-swap models — change
.ggufwithout restarting - Thinking model support —
<think>reasoning split intoreasoning_content - Chat template auto-detection — uses the template from GGUF metadata
- Zero dependencies — one
.exeon Windows, no Python, no Docker
| Platform | What you need |
|---|---|
| Windows (primary) | Windows 10/11, Go 1.25+, Vulkan-capable GPU recommended |
| Linux | Go 1.25+, curl, tar, Vulkan (Mesa/NVIDIA driver; check with vulkaninfo) or CPU |
| macOS | Go 1.25+, CPU backend (Vulkan varies by hardware) |
Disk: plan for the model size (often 2–8 GB per model) plus ~50 MB for Janus + llama.dll.
git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
.\build.ps1build.ps1 downloads pre-built llama.cpp Vulkan DLLs and compiles dist\janus.exe.
Put a .gguf file in the models\ folder. Use the included downloader:
go build -o dist\modelget.exe .\cmd\modelget
.\dist\modelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .\models\Or download any GGUF from Hugging Face.
copy .env.example .envEdit .env:
INFERENCE_BACKEND=vulkan
JANUS_MODEL_PATH=./models/Llama-3.2-3B-Instruct-Q8_0.gguf
JANUS_MAX_TOKENS=4096| Variable | Default | Meaning |
|---|---|---|
INFERENCE_BACKEND |
vulkan |
vulkan, cpu, or openrouter |
JANUS_MODEL_PATH |
(required) | Path to your .gguf file |
JANUS_GPU_LAYERS |
-1 |
-1 = all layers on GPU, 0 = CPU only |
JANUS_VRAM_CEILING_MB |
9216 |
VRAM budget hint (MiB) |
JANUS_MAX_TOKENS |
4096 |
Max tokens per reply |
JANUS_LISTEN_ADDR |
127.0.0.1:8990 |
Bind address |
.\dist\janus.exeOpens http://127.0.0.1:8990 in your browser.
curl http://127.0.0.1:8990/healthgit clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
./build.sh # downloads llama.cpp b11146 (Vulkan) into lib/linux, builds dist/janus
cp .env.example .env
# edit .env — set JANUS_MODEL_PATH (and INFERENCE_BACKEND=cpu if no Vulkan)
./dist/janusbuild.shpins llama.cpp b11146 (the struct layouts ininternal/bridge/llama_dl.gomatch that tag);./build.sh --version bNNNNtries another one,--skip-downloadreuseslib/linux/.- The
.sofiles are copied next to the binary (dist/); they find each other viaRUNPATH=$ORIGIN, noLD_LIBRARY_PATHneeded. - Go 1.25+ is required (purego v0.11 is needed to pass llama.cpp's structs by value on Linux). With
GOTOOLCHAIN=auto(default) an oldergodownloads the right toolchain itself. JANUS_GPU_LAYERS=-1(default) offloads every layer; check the log foroffloaded N/N layers to GPU.
Supported: linux/amd64; arm64/macOS not tested/not supported.
macOS: build with go build -o dist/janus ./cmd/janus and provide libllama.dylib next to the binary (untested).
curl http://127.0.0.1:8990/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local",
"messages": [{"role": "user", "content": "Hello!"}]
}'Base URL: http://127.0.0.1:8990/v1
| Method | Path | Description |
|---|---|---|
| GET | /health |
Liveness check (?deep=true for details) |
| GET | /v1/models |
Model list |
| POST | /v1/chat/completions |
Chat (streaming supported) |
| POST | /models/load |
Hot-swap model |
| GET | /models/list |
Available .gguf files |
| GET | /engine/status |
VRAM and backend info |
curl http://127.0.0.1:8990/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"local","stream":true,"messages":[{"role":"user","content":"Tell me a joke"}]}'Base URL: http://127.0.0.1:8990/v1
API Key: (leave blank)
| Problem | What's going on | Fix |
|---|---|---|
| "It built but my changes aren't there" | On Windows, Go can't overwrite a running .exe. |
Stop all janus.exe in Task Manager, then rebuild. |
| "Address already in use" | A leftover process holds port 8990. | Task Manager → end all janus.exe. |
| Server starts, chat fails | JANUS_MODEL_PATH wrong or no .gguf in models/. |
Set the path in .env, put the file in models/. |
| "local engine failed to start" | Missing llama.dll or GPU driver issue. |
Run .\build.ps1. Update GPU drivers, or set INFERENCE_BACKEND=cpu. |
| Wrong URL | Default is http://127.0.0.1:8990, not 8080. |
Bookmark 8990. |
| First reply takes forever | Model loading into VRAM — normal. | Wait 10–60s; smaller quants (Q4) load faster. |
cmd/janus/ Main server (OpenAI-compatible API)
cmd/modelget/ Hugging Face model downloader
internal/engine/ llama.cpp Vulkan/CPU backend
internal/bridge/ DLL loader and FFI bindings
internal/singleton/ Single-instance guard
models/ Put .gguf files here (not committed)
dist/ janus.exe + llama.dll after build
Get-Process -Name "janus" -ErrorAction SilentlyContinue | Stop-Process -Force
go test ./...
go build -o dist\janus.exe .\cmd\janus
.\dist\janus.exeOr just .\run.ps1. For a full build including llama DLLs, use .\build.ps1.
Contributions welcome — see CONTRIBUTING.md.
MIT — see LICENSE.
