Repository navigation
llama-modes v0.4.0
llama-modes v0.4.0
Structured LLM inference modes for llama.cpp.
This is the first public release of llama-modes.
Three structured inference modes
- BOOLEAN — direct Yes / No evaluation
- CHOICE — score arbitrary candidate labels, including multi-token choices
- SCALE — discrete ordinal or interval distributions with mode, median, quantiles, and interval statistics
Unlike normal chat generation, these modes score supplied alternatives directly at the prepared model evaluation state.
Normal llama-server chat and completion APIs remain available.
Windows CUDA runtime
This release includes a prebuilt Windows CUDA runtime.
Download: llama-modes-v0.4.0-win-cuda.zip
Bring your own GGUF model.
Validated configurations include GPT-OSS 20B MXFP4 and Qwen3.8 Ridge on Windows with NVIDIA CUDA. These tested configurations do not imply universal compatibility with every model, template, or backend.
SHA256
C5AD505CABDAAFE7D4C3A391694524DB1EFB5996939FADF2DEDF592E14D9B6B2
Quick start
.\llama-server.exe `
-m "C:\models\your-model.gguf" `
--host 127.0.0.1 `
--port 8080 `
-ngl 99 `
--jinjaStructured endpoints:
POST /decisionPOST /v1/decisionPOST /scalePOST /v1/scale
Important interpretation note
Candidate and SCALE weights are relative to the supplied alternatives and their representations.
They are not calibrated confidence.
Direct evaluation is also not equivalent to allowing a model to perform autoregressive reasoning and then answer. The two procedures may produce different results.
Release progression
- v0.1 — direct single-token decision scoring
- v0.2 — native messages and template-aware evaluation
- v0.3 — arbitrary multi-token candidate scoring
- v0.3.1 — correctness-first prompt re-prefill for multi-token scoring
- v0.4 — ordinal and interval SCALE
Public project tools
The repository also includes:
- a local React demonstration UI
- a detailed practical cookbook
- Python, PowerShell, and curl examples
- a reproducible Direct-vs-Chat benchmark harness
- documentation of design choices, limitations, and SCALE semantics
Based on llama.cpp
llama-modes is a fork/extension of ggml-org/llama.cpp.
It retains the llama.cpp runtime foundation, MIT license, attribution, and source history.