Skip to content

llama-modes v0.4.0

Choose a tag to compare

@loo5x loo5x released this 27 Sep 07:27
· 15 commits to main since this release

llama-modes v0.4.0

Structured LLM inference modes for llama.cpp.

This is the first public release of llama-modes.

Three structured inference modes

  • BOOLEAN — direct Yes / No evaluation
  • CHOICE — score arbitrary candidate labels, including multi-token choices
  • SCALE — discrete ordinal or interval distributions with mode, median, quantiles, and interval statistics

Unlike normal chat generation, these modes score supplied alternatives directly at the prepared model evaluation state.

Normal llama-server chat and completion APIs remain available.

Windows CUDA runtime

This release includes a prebuilt Windows CUDA runtime.

Download: llama-modes-v0.4.0-win-cuda.zip

Bring your own GGUF model.

Validated configurations include GPT-OSS 20B MXFP4 and Qwen3.8 Ridge on Windows with NVIDIA CUDA. These tested configurations do not imply universal compatibility with every model, template, or backend.

SHA256

C5AD505CABDAAFE7D4C3A391694524DB1EFB5996939FADF2DEDF592E14D9B6B2

Quick start

.\llama-server.exe `
  -m "C:\models\your-model.gguf" `
  --host 127.0.0.1 `
  --port 8080 `
  -ngl 99 `
  --jinja

Structured endpoints:

  • POST /decision
  • POST /v1/decision
  • POST /scale
  • POST /v1/scale

Important interpretation note

Candidate and SCALE weights are relative to the supplied alternatives and their representations.

They are not calibrated confidence.

Direct evaluation is also not equivalent to allowing a model to perform autoregressive reasoning and then answer. The two procedures may produce different results.

Release progression

  • v0.1 — direct single-token decision scoring
  • v0.2 — native messages and template-aware evaluation
  • v0.3 — arbitrary multi-token candidate scoring
  • v0.3.1 — correctness-first prompt re-prefill for multi-token scoring
  • v0.4 — ordinal and interval SCALE

Public project tools

The repository also includes:

  • a local React demonstration UI
  • a detailed practical cookbook
  • Python, PowerShell, and curl examples
  • a reproducible Direct-vs-Chat benchmark harness
  • documentation of design choices, limitations, and SCALE semantics

Based on llama.cpp

llama-modes is a fork/extension of ggml-org/llama.cpp.

It retains the llama.cpp runtime foundation, MIT license, attribution, and source history.