Skip to content
github-actions[bot] edited this page Sep 30, 2026 · 6 revisions

jevos: yes/no decisions from a local LLM on the CPU

jevos is a small local LLM that answers yes/no questions about a text. You send the text and the question, and it sends back P(yes), the probability that the answer is yes. It runs on a laptop CPU, in 26 ms for a short request and 112 ms for a 190-token one, and it generates no text, so there is nothing to parse.

That narrow job is the point. A router, a filter, a policy check, a judge in an evaluation loop: most decisions an application asks a language model for are yes/no questions with a threshold on top, and they do not need a chat model on a GPU or a call to a hosted API.

This wiki is the long version of the README: how to use the server, how to turn other kinds of decisions into yes/no questions, what we measured about the model, including where it is wrong, and how it compares with the alternatives.

jevos playing a browser game on the CPU, answering two yes/no questions per step (recording at 2x speed)

Start here

What we measured

How it compares

Browse by topic

  • Speed: where LLM latency comes from, and why a model that generates nothing is fast.
  • Probability and thresholds: what P(yes) means, calibration, and how to turn a probability into a decision.
  • Question design: how to write questions a small model answers well.
  • Use cases: moderation, triage, routing, screening, documents.
  • Evaluation: grading RAG pipelines, judges, CI checks and test sets.
  • Agents and routing: routers, cascades, tool gating and guardrails.
  • Integrations: Python, JavaScript, curl, n8n, Slack, CI and more.
  • Local and private AI: self-hosted, offline and on-premise decisions.
  • llama.cpp and GGUF: the runtime and the file format of the release's GGUF builds, and the tokenizer jev takes from llama.cpp.

What it is not

jevos answers yes/no questions, and multiple-choice questions asked as one yes/no question per option, in English only. Scores, asked as one yes/no question per level, are early (54% on held-out score questions, 82% within one level). On rules it has never seen it is right about four times in five (0.810 on 2,000 such questions), which is good for a first pass and not good enough to be the last word on a refund, and the measurement pages say exactly where it fails.


jevos is built by feder-cr with Loris Salsi (@LosaLosSantos). The code is MIT, the model is on the release page.

Guides

Measurements

Comparisons

Speed

Probability and thresholds

Question design

Use cases

Evaluation

Agents and routing

Integrations

Local and private AI

llama.cpp and GGUF

Clone this wiki locally