-
Notifications
You must be signed in to change notification settings - Fork 129
evaluation metrics for yes no classifiers
For a yes/no classifier that returns a probability, report at least four things: accuracy per kind of question, the split between wrong yeses and wrong noes, precision and recall at the threshold you will actually use, and a calibration measure. Each one answers a different question, and each one can look fine while another is bad. Add an uncertainty range to every number computed on fewer than a few hundred cases.
The trap is the single accuracy figure. On our 999-question test set jevos scored 0.757 overall, which describes no real use: the same model was right 0.954 of the time on stated facts and 0.584 on arithmetic, and when it was wrong it was wrong toward yes five times out of eight.
This page is the confusion matrix every metric comes from, a table of which metric answers which question, the three reports people most often skip, threshold-free scores, and how much a number on a small set can be trusted.
With a threshold (say P(yes) above 0.5 means yes), every answer falls in one of four cells. An illustrative example, not a measurement, for 200 cases with 100 of each answer:
| answer is yes | answer is no | |
|---|---|---|
| model says yes | 85 (true yes) | 30 (wrong yes) |
| model says no | 15 (wrong no) | 70 (true no) |
From these four numbers:
- accuracy = (85 + 70) / 200 = 0.775: share of answers that are right.
- precision = 85 / (85 + 30) = 0.74: of the yeses, how many were right.
- recall = 85 / (85 + 15) = 0.85: of the real yeses, how many were found.
- specificity = 70 / (70 + 30) = 0.70: of the real noes, how many were kept as no.
Move the threshold and all four cells change. That trade is worth a page of its own: precision and recall at a P(yes) threshold.
| You want to know | Report |
|---|---|
| how often is it right overall, on this mix | accuracy |
| when it says yes, can I act on it | precision |
| does it find the cases I care about | recall |
| which kinds of question can I trust it with | accuracy per kind |
| which way does it fail | wrong yes vs wrong no counts |
| can I read 0.8 as "about 80% likely" | calibration error, reliability diagram |
| how good is it before I pick a threshold | log loss, Brier score, ROC AUC |
| how much would the number move on another sample | a confidence interval |
Tag every test question with the reasoning it needs (stated fact, tone, negation, rule, number, date, arithmetic) and report accuracy for each tag. On our set the spread runs from 0.954 to 0.584, and the line between reading and computing is sharp. That split decides what you route to the model and what you keep in code, which one overall number never can. How to design the tags is on accuracy by kind of question.
Accuracy treats a wrong yes and a wrong no as equal. Applications almost never do: a wrong yes might refund a fraudster, a wrong no might ignore an urgent ticket. So count them separately.
On our 999 questions, half yes and half no, jevos made 152 wrong yeses and 91 wrong noes. On a balanced set, an unbiased model would split its errors roughly evenly. A clearer view is the mean P(yes) on questions whose answer is no, per kind: 0.59 on arithmetic, 0.16 on tone. The details are on why a small LLM says yes. If you report only accuracy, this is invisible.
A calibrated model's answers of about 0.8 are right about 80% of the time. The usual summary is the expected calibration error (ECE), the weighted average gap between confidence and accuracy across bins. On the natural yes/no questions of our held-out split jevos had an ECE of 0.009, and that number did not predict the lean toward yes on new kinds of question, which is why calibration belongs next to per-kind accuracy and not in place of it. The definition and a worked example are on expected calibration error, explained; drawing the picture is on reading a reliability diagram.
When you have not picked a threshold yet, or want to compare two models independently of one, use the probabilities directly:
- Log loss: the mean of minus the log of the probability given to the right answer. It punishes confident mistakes hard: on a case whose answer is no, a P(yes) of 0.99 costs about 4.6, a P(yes) of 0.6 about 0.9, and a P(yes) of 0.1 about 0.1.
- Brier score: the mean squared gap between the probability and the answer (1 or 0). Easier to read than log loss, gentler on confident mistakes.
- ROC AUC: the probability that a random yes case gets a higher P(yes) than a random no case. It measures ranking only, so a model can have a good AUC and be badly calibrated.
These are good for comparing builds or prompts. They are poor for explaining results to people who will act on thresholds; there, the confusion matrix at the chosen threshold is clearer.
Less than it looks. A rough 95% interval for an accuracy p on n cases is plus or minus
1.96 * sqrt(p * (1 - p) / n).
- 0.8 on 100 cases: about plus or minus 0.08.
- our arithmetic kind, 0.584 on 221 questions: about plus or minus 0.065.
- our tone kind, 0.938 on 32 questions: about plus or minus 0.08, and the formula is optimistic this close to 1.
So a per-kind accuracy on 30 questions tells you "high" or "low", not the second decimal, and a difference of a few points between two prompts on 100 cases is usually noise. More cases per kind is the fix; generated questions with computed answers are the cheap way to get them, described on generating test questions with answers computed by code.
What metrics should I use for a binary classifier? Accuracy per kind, error direction, precision and recall at your threshold, and a calibration measure, each with an uncertainty range.
Is accuracy enough? No. Our overall 0.757 hid a range from 0.954 to 0.584 and a lean toward wrong yeses.
What is the difference between precision and recall? Precision is the share of predicted yeses that are right; recall is the share of real yeses that were found.
When should I use log loss or AUC? To compare models or prompts before choosing a threshold. Once a threshold is set, report the confusion matrix at it.
How many test cases do I need? Enough per kind that the interval is narrower than the differences you care about. About 100 gives plus or minus 8 points.
See also: building a yes/no test set for your own data, LLM calibration explained with yes/no answers and our held-out benchmark said 0.855, new questions said 0.757.
- Overall and per-kind accuracy, error counts and mean P(yes) on no-answer questions: our
999-question test set,
jevos-q4_k_m. - Calibration error 0.009: our held-out split, 6,397 natural yes/no questions,
jevos-q8_0. - The confusion matrix and the log-loss examples are illustrative arithmetic, not measurements. The interval formula is the standard normal approximation for a proportion.
From the notes of jev, which returns a probability rather than a label, so every metric on this page can be computed from its answers.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building