-
Notifications
You must be signed in to change notification settings - Fork 132
why wording changes the answer
A language model answers the words you send, not the meaning you intended, so two phrasings that mean the same thing to you can get different probabilities. The practical fix is a small test: write three phrasings of the question, run them on 30 to 50 cases whose answer you know, and keep the phrasing that is right most often and moves least between similar cases. It takes an afternoon and replaces a guess with a measurement.
Wording sensitivity is not a flaw of small models only. Sclar and colleagues measured differences of up to 76 accuracy points on one 13B model from formatting changes alone, and found the sensitivity did not go away with larger models, more examples or instruction tuning. If formatting moves a large model that much, the choice of words in a yes/no question deserves a test on any model.
This page is where the sensitivity comes from, which changes tend to matter, the test itself, and how to read its results. It contains no measurement of jevos on paraphrased questions, because we have not run one; it tells you how to run yours.
A question is not a query into a database. The model reads it together with the text and produces a probability from both, and every word shifts that reading a little. Most shifts are small. Some are not, and they tend to come from a few sources:
- Which words overlap with the text. "Does the customer mention a refund?" and "Does the customer ask for their money back?" point at different words. If the message says "I want my money back", the second question lines up with it more directly.
- How strict the verb is. "Mention", "ask for", "demand" and "suggest" set different bars. The model will apply whichever bar you wrote, and you may have meant another.
- Who the subject is. "Is this rude?" leaves open whether you mean the customer, the agent, or the situation.
- Hidden extra conditions. "Is the customer clearly upset?" adds "clearly", which is a second condition about strength.
None of this is special to machines. A person given these questions would also answer them differently. The difference is that a person asks what you meant, and a model does not.
1. Collect labelled cases. 30 to 50 real texts, with the answer to your question decided by a person who knows what the question is for. Aim for roughly as many yes as no, and include the hard ones: short messages, sarcasm, messages about two things. How to build a set that tells you something is on building a yes/no test set for your own data.
2. Write three phrasings. Vary one thing at a time where you can: the verb, the subject, the presence of a qualifier. For example:
"questions": {
"a": {"type": "noul", "instructions": "Does the customer ask for a refund?"},
"b": {"type": "noul", "instructions": "Does the customer ask to get their money back?"},
"c": {"type": "noul", "instructions": "Is the customer requesting a refund for this order?"}
}All three can go in one request. They share the reading of the text, so testing three phrasings costs far less than three separate runs.
3. Score each phrasing. For every phrasing, record accuracy at your threshold and the probability on each case. Keep the numbers per case, not only the average.
# probs[i][name]: the noul for phrasing `name` on case i; labels[i]: True or False
for name in ["a", "b", "c"]:
right = sum((probs[i][name] > 0.5) == labels[i] for i in range(len(labels)))
print(name, right / len(labels))Look at three things, in this order:
- Accuracy. The obvious one. With 40 cases, one case is 2.5 points, so do not pick a winner on a one-case difference.
- Agreement between phrasings. Cases where all three agree are cases the model reads the same way whatever you write; they are safe. Cases where they disagree are where your wording decides the answer. Read those texts: they usually show which phrasing matches what you meant.
- Margin. A phrasing whose correct answers sit at 0.9 and 0.1 is more robust than one whose correct answers sit at 0.6 and 0.4, even at the same accuracy. Small wording changes later, or new kinds of text, will flip the second one first.
If one phrasing wins on all three, use it. If they are all similar, pick the plainest one, since it is the one colleagues will not "improve" by accident later.
Before testing, apply the rules that remove the common causes: one condition per question, the subject named, the rule written in, and computation moved to code. They are collected on how to write yes/no questions an LLM answers well. Two more are specific to wording:
- Use the words the texts use. If customers write "money back", a question with "money back" reads more directly than one with "reimbursement".
- Drop qualifiers you would not act on. "Clearly", "really", "strongly" raise the bar in a way you cannot see. If you need levels of strength, ask them as boundaries, as on scores as yes/no thresholds.
The labelled set and the three phrasings are worth keeping. When you change the question later, or move to a new model file, rerun them. That turns the set into a regression test, which is the subject of LLM regression tests in CI with yes/no checks.
Being straight about the limit: a 40-case test tells you which phrasing is better on those 40 cases. It will not tell you the accuracy you will see in production to two decimals. Use it to choose, not to report.
Why does rephrasing a prompt change the answer? Every word shifts how the model reads the question against the text. Verbs, subjects and qualifiers change the bar the answer has to meet.
How sensitive are LLMs to prompt wording? One study measured up to 76 accuracy points of difference from formatting alone on a 13B model. We have not measured jevos on paraphrases.
How many phrasings should I test? Three is enough to see whether the answer depends on wording. More rarely changes the choice.
How many labelled cases do I need? 30 to 50 to choose between phrasings. More if the decision is costly or the texts vary a lot.
Should I average the three phrasings instead of picking one? It can smooth the answer, but it triples the questions per request. Pick one unless the phrasings disagree often and you cannot tell which is right.
See also: accuracy by kind of question, negation in yes/no questions and LLM judge bias and how to control it.
- Melanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design", 2023, arXiv:2310.11324, fetched 2026-09-29.
- No jevos measurement is quoted on this page. The test described is a method, not a result.
From the notes of jev. We have not measured jevos on paraphrased questions, so this page gives you the test instead of a number.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building