-
Notifications
You must be signed in to change notification settings - Fork 129
Home
jevos is a small local LLM that answers yes/no questions about a text. You send the text and the question, and it sends back P(yes), the probability that the answer is yes. It runs on a laptop CPU, in 26 ms for a short request and 112 ms for a 190-token one, and it generates no text, so there is nothing to parse.
That narrow job is the point. A router, a filter, a policy check, a judge in an evaluation loop: most decisions an application asks a language model for are yes/no questions with a threshold on top, and they do not need a chat model on a GPU or a call to a hosted API.
This wiki is the long version of the README: how to use the server, how to turn other kinds of decisions into yes/no questions, what we measured about the model, including where it is wrong, and how it compares with the alternatives.

- Ask a local LLM a yes/no question and get P(yes): the request, the answer, Python, the command line, and what to do with the number.
- Zero-shot text classification with yes/no questions: routing a ticket to one of several labels with one question per label.
- LLM policy decisions: put the rule in the question: refunds, access rules and thresholds, and why rules are the model's hardest case.
- LLM as a judge on a CPU: evaluation criteria as yes/no questions, scored locally.
- Why a small LLM says yes when the answer is no: 152 wrong yeses against 91 wrong noes on 999 new questions, and why calibration does not fix it.
- Small LLMs and arithmetic in yes/no questions: 0.58 on questions that need a sum, against 0.95 on questions that only need reading.
- Our held-out benchmark said 0.855, new questions said 0.757: why a test split built like the training data is not held out enough.
- jevos vs Jev vs Laya for yes/no decisions: speed, accuracy, context, cost and what each one can answer.
- Speed: where LLM latency comes from, and why a model that generates nothing is fast.
- Probability and thresholds: what P(yes) means, calibration, and how to turn a probability into a decision.
- Question design: how to write questions a small model answers well.
- Use cases: moderation, triage, routing, screening, documents.
- Evaluation: grading RAG pipelines, judges, CI checks and test sets.
- Agents and routing: routers, cascades, tool gating and guardrails.
- Integrations: Python, JavaScript, curl, n8n, Slack, CI and more.
- Local and private AI: self-hosted, offline and on-premise decisions.
- llama.cpp and GGUF: the runtime and the file format of the release's GGUF builds, and the tokenizer jev takes from llama.cpp.
jevos answers yes/no questions, and multiple-choice questions asked as one yes/no question per option, in English only. Scores, asked as one yes/no question per level, are early (54% on held-out score questions, 82% within one level). On rules it has never seen it is right about four times in five (0.810 on 2,000 such questions), which is good for a first pass and not good enough to be the last word on a refund, and the measurement pages say exactly where it fails.
jevos is built by feder-cr with Loris Salsi (@LosaLosSantos). The code is MIT, the model is on the release page.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building