-
Notifications
You must be signed in to change notification settings - Fork 130
reading a reliability diagram
A reliability diagram plots, for bins of predicted P(yes), the fraction of cases that were really yes against the mean prediction in the bin; a calibrated model sits on the diagonal. A curve flatter than the diagonal means the model is overconfident, a steeper one means it is underconfident, and a curve that sits below the diagonal all along means it says yes too readily. You draw it from your own labelled answers, and you read it together with the number of answers in each bin, because a point made of eight cases is mostly noise.
It is the most useful single picture of a probability model, because it shows where the numbers are wrong, not only that they are. One number such as ECE averages the bins away; the diagram keeps them.
This page is how to draw one from your answers, the shapes to recognise, how many answers each bin needs, why to draw one per kind of question, and what the diagram does not show.
- Collect labelled cases. A few hundred real texts with the questions you will use in production, each answered yes or no by a person who read the text.
- Run them and keep every P(yes). One request per text, one named question per condition.
-
Bin. scikit-learn's
calibration_curvedoes this: it splits [0, 1] inton_binsbins (default 5) and returns, per bin, the mean predicted probability and the fraction of positives. Bins with no samples are dropped.strategy="uniform"gives equal widths;strategy="quantile"gives each bin the same number of samples. - Plot the mean prediction on the x-axis and the fraction of yes on the y-axis, with the diagonal from (0, 0) to (1, 1) for reference, and a histogram of the counts underneath.
from sklearn.calibration import calibration_curve
frac_yes, mean_p = calibration_curve(y_true, p_yes, n_bins=10, strategy="quantile")y_true is 1 where the true answer is yes; p_yes holds the probabilities your requests
returned.
| What you see | What it means | What usually helps |
|---|---|---|
| points on the diagonal | calibrated on this data | nothing; keep checking |
| curve flatter than the diagonal (low bins above it, high bins below) | overconfident: numbers too close to 0 and 1 | a temperature above 1 |
| curve steeper than the diagonal (low bins below, high bins above) | underconfident: numbers too close to 0.5 | a temperature below 1 |
| curve below the diagonal across the range | says yes too readily at every level | a higher yes threshold, a Platt offset, or better questions |
The last row is worth a second look. Below the diagonal means that among the answers with a given P(yes), fewer were yes than claimed. If only a few bins sit below, and they are made of one kind of question, the fix is the question, not a curve. The two corrections in the right-hand column are on temperature scaling and Platt scaling.
Published diagrams make the shapes concrete. The GPT-4 technical report shows two, for the pre-trained and the post-trained model on a subset of MMLU, with ECEs of 0.007 and 0.074; the second is visibly further from the diagonal.
Guo and colleagues (2017) point out that a reliability diagram does not show how many samples are in each bin, which hides whether a point is solid evidence. The arithmetic of a proportion says how much to trust it. With n answers in a bin and a true frequency near p, the standard error of the observed fraction is about the square root of p(1 - p) / n. For p = 0.8:
| Answers in the bin | Standard error | A point at 0.8 could plausibly be |
|---|---|---|
| 25 | 0.08 | anywhere from about 0.64 to 0.96 |
| 100 | 0.04 | about 0.72 to 0.88 |
| 400 | 0.02 | about 0.76 to 0.84 |
This is textbook arithmetic, not a measurement. It means a 10-point gap in a bin of 25 answers is not yet evidence of anything, while the same gap in a bin of 400 is. Practical rules:
- Use quantile bins when your probabilities cluster near 0 and 1, which they usually do; equal widths leave the middle bins nearly empty.
- With a few hundred answers, use 5 to 10 bins, not 15.
- Always print the count per bin next to the plot.
On the natural yes/no questions of our held-out split, jevos-q8_0 had a calibration error of
0.009, which means a pooled diagram on that data hugs the diagonal. On 999 new questions,
labelled by the kind of reasoning they need, the mean P(yes) on questions whose answer is no
ranged from 0.16 for tone to 0.59 for arithmetic. A pooled diagram of that set mixes a
well-behaved kind with a leaning one. Split by kind, you would expect the arithmetic and date
curves to sit below the diagonal, since their no-answers get an average P(yes) of 0.59 and
0.53, while tone and negation stay much closer to it. The per-kind figures are on
why a small LLM says yes when the answer is no, and how to tag
your own questions by kind is on accuracy by kind of question.
So tag each labelled case with the kind of question (reading a fact, tone, a date, a sum, a rule) and draw a diagram per tag once each has enough answers.
- Whether the model separates yes from no. A model that returns 0.5 to everything on a half-yes set produces one point, on the diagonal. It is calibrated and useless. Check discrimination with precision and recall at a threshold.
- Errors that cancel inside a bin. Two opposite mistakes averaged together look like a perfect point.
- Anything about data you did not label. The diagram describes the mix of cases you drew it from. If production traffic differs, draw it again on production traffic.
What is a reliability diagram? A plot of observed frequency against predicted probability, bin by bin, with the diagonal as the calibrated reference.
What does a point below the diagonal mean? Among answers with that P(yes), fewer were yes than predicted: the model overstates yes there.
How many bins should I use? With a few hundred labelled answers, 5 to 10, preferably with equal counts per bin.
Is a calibration curve the same as a reliability diagram? Yes, the two names are used for the same plot; scikit-learn calls it a calibration curve.
Can a model be on the diagonal and still bad? Yes. Calibration says the numbers are honest, not that they are sharp.
See also: expected calibration error, explained, LLM calibration explained and building a yes/no test set.
- Our measurements: calibration error 0.009 on 6,397 natural yes/no held-out questions
(
jevos-q8_0); mean P(yes) on no-answer questions by kind, 999-question set (jevos-q4_k_m). - The standard-error table is arithmetic for illustration, not a measurement.
- scikit-learn calibration_curve and Probability calibration, fetched 2026-09-29.
- Guo, Pleiss, Sun, Weinberger (2017), On Calibration of Modern Neural Networks, fetched 2026-09-29.
- OpenAI (2023), GPT-4 Technical Report, Figure 8, fetched 2026-09-29.
From the notes of jev, where a pooled calibration error of 0.009 did not predict the yes-lean on new kinds of question, which is why this page insists on splitting by kind.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building