-
Notifications
You must be signed in to change notification settings - Fork 132
llm judge bias
LLM judges have systematic biases: they favour an answer because of its position, because it is longer, because the question invites a yes, or because it is written in their own style, and their verdict moves when the criterion is reworded. None of these shows up in a single verdict, and all of them show up in a set of cases whose right answers you know. The controls are correspondingly simple: swap the order, compare equal-length outputs, measure the judge on a labelled set per criterion, and look at the direction of its errors, not only their number.
The bias that matters most for a yes/no judge is the one people least expect: a lean toward yes. It is invisible in accuracy and calibration measured on familiar data, and it inflates every pass rate you compute.
This page is a map of the biases with a detection test for each, what we measured on jevos, the known-answer set that catches most of them, and what a judge bias does to the numbers you report.
| Bias | What it looks like | Detection test | Control |
|---|---|---|---|
| position | prefers the first (or second) of two answers | ask both orders; count verdicts that follow the position | swap and average; ties when orders disagree |
| length | prefers the longer answer | pairs that differ only by padding or repetition | compare similar lengths; add "repeats itself?" |
| yes-lean | says yes when unsure | mean P(yes) on cases whose answer is no | higher bar for yes; keep computing out of the judge |
| self-preference | prefers text in its own style | compare verdicts on outputs from different generators with the same labels | use a judge from another family; check with labels |
| wording | verdict moves with the phrasing | two phrasings of one criterion on the same cases | pick the stable phrasing; freeze it |
Zheng and colleagues, in "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", name position, verbosity and self-enhancement biases, together with limited reasoning ability, as the main limitations of LLM judges. The yes-lean and the wording effect are the two we add for yes/no judges in particular.
Position bias only exists when the judge sees two outputs at once, and the fix is procedural: ask in both orders. The formula and the tie rule are on pairwise comparison with a yes/no judge.
Length bias survives the swap, because the longer answer is longer in both orders. In Zheng and colleagues' test, where answers were made longer by repeating content, Claude-v1 and GPT-3.5 each failed 91.3% of the attacks and GPT-4 failed 8.7%. For a yes/no judge, the practical controls are to compare outputs of similar length, to ask pointwise questions ("Does the reply answer the question in its first two sentences?") rather than "which is better", and to add a negative criterion for padding.
On 999 yes/no questions written after training, exactly half with the answer yes, jevos made 152 mistakes by saying yes when the answer was no and 91 the other way. Accuracy (0.757) says nothing about the direction; the error split does.
The better diagnostic is the mean P(yes) on questions whose right answer is no. For an ideal judge it is 0. For jevos it ranged from 0.16 on tone questions to 0.59 on arithmetic. So the lean is not a general optimism: it appears where the judge cannot work the answer out and the question reads like a yes. The full breakdown, and why recalibrating did not remove it, is on why a small LLM says yes when the answer is no.
For evaluation the consequence is direct: a judge that leans toward yes reports inflated pass rates, most of all on the criteria it is worst at. Two controls:
- Keep computing out of the judge. Criteria about sums, dates and thresholds go to code.
- Raise the pass bar on criteria where a false pass is costly. How far to raise it is a cost question, covered on thresholds when a wrong yes costs more than a wrong no.
A judge from the same family as the generator may rate that generator's style higher; Zheng and colleagues call this self-enhancement bias. jevos generates no text, so it is never judging its own writing. That is not the same as having no style preference: a judge can still favour outputs that resemble text it saw often. The only way to know is the known-answer test below, run separately on outputs from each generator you compare.
The same criterion phrased two ways can give different pass rates on the same outputs. This is less a bias of the judge than a sign that the criterion is ambiguous, and it is cheap to test: run two or three phrasings on a labelled sample and keep the one that agrees best with the labels. The method is on why wording changes an LLM's answer.
Every bias above is detectable with the same tool: a set of cases where you know the right answer for each criterion.
- Take 50 to 100 real outputs, labelled by people, per criterion. Include outputs from each generator you plan to compare, and pairs of similar and different lengths.
- Run the judge on them in both orders where there are pairs.
- Report, per criterion: accuracy, the split between wrong yes and wrong no, mean P(yes) on the no cases, and for pairs the share of order-dependent verdicts.
- Re-run it whenever the judge, the rubric or the generator changes.
A criterion whose wrong answers are mostly in one direction needs a different threshold. A criterion where the judge is near chance needs to leave the rubric. How to assemble the set, and keep it clean, is on building a yes/no test set.
A judge's error does not average out when you compare systems. If it leans toward yes on a criterion, both systems get inflated pass rates on it, and the difference between them can still be real. But when one system produces the kind of output the judge is biased toward (longer, first, in a familiar style), the bias becomes part of the measured difference. So report per criterion, state the judge's own accuracy and error direction next to the results, and treat any difference smaller than the judge's disagreement with your labels as noise.
What is LLM judge bias? A systematic preference of an LLM judge that is not about the quality being judged: position, length, a lean toward yes, its own style, or the wording of the criterion.
Does a yes/no judge have a bias toward yes? jevos does on questions it cannot work out: 152 wrong yeses against 91 wrong noes on a balanced 999-question set.
How do I detect bias in my judge? Run it on cases with known answers and look at the direction of its errors per criterion, not only its accuracy.
Does calibration remove bias? Not a directional one. A temperature and bias fitted on our development split moved accuracy from 0.757 to 0.759.
Is a judge from a different model family unbiased? No, only free of self-preference toward the generator. It still has position, length and yes-lean effects to check.
See also: LLM as a judge on a CPU, LLM calibration explained with yes/no answers and evaluation metrics for yes/no classifiers.
- Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", arXiv:2306.05685 and HTML version, fetched 2026-09-29: the named biases, the repetitive-list attack and its failure rates.
- Error split, mean P(yes) on no-answer questions and the recalibration result: our 999-question
test set,
jevos-q4_k_m.
From the notes of jev, a yes/no model whose largest measured bias is toward yes, published so you can set your bars with it in mind.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building