Skip to content

Releases: loo5x/llama-modes

llama-modes v0.5.0

Choose a tag to compare

@loo5x loo5x released this 03 Oct 18:17

llama-modes v0.5.0 — One Context, Many Questions

This release adds shared-context multi-question evaluation.

Load one GGUF model once, evaluate one shared context once, then ask multiple structured questions over that same prepared context.

Questions can mix:

  • BOOLEAN
  • CHOICE
  • SCALE

in the same request.

Why this matters

Before v0.5, applications that wanted several structured judgments over the same long input had to repeat prompt/context evaluation for each question.

v0.5 introduces a shared-context path so the common context is evaluated once and reused across multiple independent structured questions.

Conceptually:

shared context
      |
      +--> Boolean question
      +--> Choice question
      +--> Scale question
      +--> another Boolean question

The goal is to reduce repeated context processing while preserving the semantics of the existing v0.4 modes.

Correctness-first design

The shared-context implementation was developed against an independent fresh-evaluation oracle.

Validation focuses on more than the final selected answer:

  • full score-vector parity
  • original vs reversed question order
  • candidate-order invariance
  • repeated runs
  • mixed Boolean / Choice / Scale requests
  • preservation of existing scoring semantics

The optimized shared path is required to reproduce the same evaluation state and scoring behavior as independent evaluation within the expected numeric tolerance.

This follows the same correctness-first approach used after the v0.3.1 re-prefill fix.

Existing modes remain available

Normal llama-server chat/completion behavior is unchanged.

The existing structured endpoints remain available for single-question use:

  • POST /decision
  • POST /v1/decision
  • POST /scale
  • POST /v1/scale

v0.5 adds shared-context multi-question evaluation on top of those existing primitives rather than replacing them.

Interpretation

The same caveats from earlier releases still apply:

  • Direct structured inference is a readout from a particular prepared model state, not shortened Chat.
  • Candidate and SCALE weights are relative to the supplied representations.
  • They are not calibrated correctness probabilities.
  • Label wording, tokenization, chat templates, and answer/readout boundaries can affect results.
  • Direct evaluation does not universally replace model reasoning.

Demo

The repository also includes a local React demo for Shared Context, with mixed Boolean, Choice, and Scale questions over one input.

Windows CUDA runtime

This release includes a prebuilt Windows CUDA runtime.

Download:

llama-modes-v0.5.0-win-cuda.zip

Bring your own GGUF model.

SHA256

18F9FE09ECEB93D64212D89C6ECD7D0CE782223662B97F010C2A2F20B8B5342F

A matching .sha256 file is also attached to the release.

Release progression

  • v0.1 — direct single-token decision scoring
  • v0.2 — native messages and template-aware evaluation
  • v0.3 — arbitrary multi-token candidate scoring
  • v0.3.1 — correctness-first prompt re-prefill
  • v0.4 — ordinal and interval SCALE
  • v0.5 — shared-context multi-question evaluation

Based on llama.cpp

llama-modes is a fork/extension of ggml-org/llama.cpp.

It retains the llama.cpp runtime foundation, MIT license, attribution, and source history.

llama-modes v0.4.0

Choose a tag to compare

@loo5x loo5x released this 27 Sep 07:27

llama-modes v0.4.0

Structured LLM inference modes for llama.cpp.

This is the first public release of llama-modes.

Three structured inference modes

  • BOOLEAN — direct Yes / No evaluation
  • CHOICE — score arbitrary candidate labels, including multi-token choices
  • SCALE — discrete ordinal or interval distributions with mode, median, quantiles, and interval statistics

Unlike normal chat generation, these modes score supplied alternatives directly at the prepared model evaluation state.

Normal llama-server chat and completion APIs remain available.

Windows CUDA runtime

This release includes a prebuilt Windows CUDA runtime.

Download: llama-modes-v0.4.0-win-cuda.zip

Bring your own GGUF model.

Validated configurations include GPT-OSS 20B MXFP4 and Qwen3.8 Ridge on Windows with NVIDIA CUDA. These tested configurations do not imply universal compatibility with every model, template, or backend.

SHA256

C5AD505CABDAAFE7D4C3A391694524DB1EFB5996939FADF2DEDF592E14D9B6B2

Quick start

.\llama-server.exe `
  -m "C:\models\your-model.gguf" `
  --host 127.0.0.1 `
  --port 8080 `
  -ngl 99 `
  --jinja

Structured endpoints:

  • POST /decision
  • POST /v1/decision
  • POST /scale
  • POST /v1/scale

Important interpretation note

Candidate and SCALE weights are relative to the supplied alternatives and their representations.

They are not calibrated confidence.

Direct evaluation is also not equivalent to allowing a model to perform autoregressive reasoning and then answer. The two procedures may produce different results.

Release progression

  • v0.1 — direct single-token decision scoring
  • v0.2 — native messages and template-aware evaluation
  • v0.3 — arbitrary multi-token candidate scoring
  • v0.3.1 — correctness-first prompt re-prefill for multi-token scoring
  • v0.4 — ordinal and interval SCALE

Public project tools

The repository also includes:

  • a local React demonstration UI
  • a detailed practical cookbook
  • Python, PowerShell, and curl examples
  • a reproducible Direct-vs-Chat benchmark harness
  • documentation of design choices, limitations, and SCALE semantics

Based on llama.cpp

llama-modes is a fork/extension of ggml-org/llama.cpp.

It retains the llama.cpp runtime foundation, MIT license, attribution, and source history.