-
Notifications
You must be signed in to change notification settings - Fork 130
one condition per question
If a yes/no question contains "and" or "or" between two conditions, split it into one question per condition and combine the probabilities in your code. "Is the customer upset and asking for a refund?" returns a single number, and when that number is low you cannot tell whether the model found no anger, no refund request, or both. Two questions return two numbers, each of which means one thing, and the combination is a line of arithmetic you control.
Survey designers have a name for the compound form: a double-barreled question, one that touches more than one issue but allows only one answer. The problem is the same with a model as with a person. The answer you get back is to a question you cannot see, the one the reader decided to answer.
This page is how to spot a compound question, how to split it, how to put the pieces back together, and what the split costs in latency.
The obvious sign is a conjunction between conditions. The less obvious ones:
- A rule with several clauses. "Refund if the item was reported missing within 30 days and the order was over 20 dollars" is three conditions in one sentence.
- Two subjects. "Were the customer and the agent both polite?" is two questions about two people.
- A condition hidden in a noun. "Is this a repeat complaint about a late delivery?" asks whether it is a complaint, whether it is about a delivery, whether the delivery was late, and whether it has happened before.
- "Or" between labels. "Is this about billing or shipping?" is fine as a routing question only if you do not care which; usually you do.
The Wikipedia examples of double-barreled questions read like support-ticket questions: "How satisfied are you with your pay and job conditions?" and "Is this tool interesting and useful?". Both combine two things a person could feel differently about.
Take the refund rule above. As one question:
{"refund": {"type": "noul", "instructions": "Should the customer get a refund if the item was reported missing within 30 days and the order was over 20 dollars?"}}As single conditions, with the numbers left to code:
{
"model": "jev-latest",
"state": {
"order_total": "34 dollars",
"delivered": "5 days ago",
"customer_message": "The box arrived empty. This is the second time!"
},
"questions": {
"missing": {"type": "noul", "instructions": "Does the customer say an item was missing from the delivery?"},
"repeat": {"type": "noul", "instructions": "Does the customer say this has happened before?"},
"upset": {"type": "noul", "instructions": "Is the customer upset?"}
}
}The 30-day window and the 20-dollar floor are no longer questions at all. Your system has the delivery date and the order total as fields, and code compares them exactly. That matters because a rule question is where a small model is weaker: 0.721 on applying a stated rule on our 999-question set, against 0.954 on reading a stated fact, and 0.584 when the question needs arithmetic. Splitting moves each condition toward the reading end of that range. The wider version of this argument is on LLM policy decisions: put the rule in the question.
Each probability is independent, so you choose how to combine them. The three usual rules:
p = {k: a["noul"] for k, a in answers.items()}
strict = min(p["missing"], p["upset"]) # AND, cautious
joint = p["missing"] * p["upset"] # AND, if the conditions are independent
escalate = max(p["upset"], p["repeat"]) # OR, cautiousmin is the safe default for AND: the combination is only as sure as its least sure part. The
product assumes the two conditions are independent, which they often are not (an upset
customer is more likely to be reporting a problem), and it drifts low when you chain several.
Which one to use, and how NOT fits in, is the subject of
combining yes/no answers with AND, OR and NOT.
Keep the individual probabilities in your logs next to the combined decision. When a case is decided wrong, they show which condition caused it, which is the whole point of the split.
Less than you would think. Questions in the same request share the state, and the text is
read once. On the reference laptop (Intel Core Ultra 7 255H, 16 threads), the
README's three questions about one text take about 66 ms together, against 49 ms for one of
them alone. Three conditions for about 1.3 times the price of one is a trade worth making for
any decision you will need to explain. Why the extra questions are cheap is on
many questions about one text.
Not every "and" is two conditions. "Did the customer thank the agent and say goodbye?" might be a single thing you care about, a polite close, and splitting it gains nothing if you would never act on the parts separately. The test is simple: if the two parts could have different answers, and you would do something different depending on which one failed, split. If not, a single question with a clear name is fine.
The same test applies to "or". "Does the customer mention a refund or a replacement?" is one condition if both lead to the same queue.
What is a double-barreled question? A question that touches more than one issue but allows only one answer. With a yes/no model, it returns one probability for two conditions.
Should I use "and" in an LLM prompt? In a yes/no question, only when the two parts are one thing you care about. Otherwise ask them separately.
How do I combine two yes/no probabilities? For AND, the minimum is the cautious choice and the product assumes independence. For OR, the maximum. Keep the parts in your logs.
Does asking more questions make the request slow? Less than linearly. On the README example, three questions take about 66 ms against 49 ms for one, because the text is read once.
See also: how to write yes/no questions an LLM answers well, negation in yes/no questions and rubric design for an LLM judge.
- Accuracy on rule, fact and arithmetic questions: our 999-question test set,
jevos-q4_k_m. - The 66 ms and 49 ms timings: the jev README.
- Definition and examples of double-barreled questions: Double-barreled question on Wikipedia, fetched 2026-09-29.
From the notes of jev, a model that returns exactly one probability per question, which is the reason each question should mean exactly one thing.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building