-
Notifications
You must be signed in to change notification settings - Fork 131
generating test questions with code
A generated test set is a small program that writes texts from templates, asks yes/no questions about them, and computes each right answer from the same values it put in the text. Because code decides the answer, there are no labelling mistakes, the set can be regenerated at any size with a fixed seed, and every question kind can have as many cases as you need. The price is a style of its own: templated texts are cleaner than real ones, so a generated set is a complement to a hand-labelled set of real cases, not a replacement.
It is most valuable exactly where hand labelling is weakest: questions with numbers, dates and rules, where a tired labeller adds wrong, and where you need hundreds of cases per kind before the per-kind accuracy means anything.
This page is what a generator is good for, a minimal generator in Python, how to control balance, boundaries and phrasing, how to keep it reproducible, and the biases a generated set brings with it.
- Correct answers. "Do the three items cost more than 110 euros together?" has one answer, and the generator knows it because it chose the prices.
- Size on demand. A kind with 30 hand-written questions has a range of about plus or minus 8 points. A generator gives you 500 of that kind in a second.
- Control. You decide how many cases sit near the threshold, how many are far from it, and how the yeses and noes are balanced.
- Regeneration. When a test set has been looked at too often while tuning, a new seed gives a fresh one of the same shape.
It suits question kinds whose answer follows from structured values: number against a threshold, arithmetic, dates and durations, rules with explicit conditions, and "is it stated" questions where the generator decides which fields to include. It does not suit tone, intent or anything where a person's judgement is the definition of the right answer.
This sketch writes request files in the documented POST /v1/systemone body, plus an expectation
file per case in the layout used on
LLM regression tests in CI. It is written for this page, not a
tested tool.
import json, pathlib, random
rng = random.Random(7) # fixed seed: same set every run
TEMPLATES = [
"Is the order total over {t} euros?",
"Does the order cost more than {t} euros in total?",
"Taken together, do the items come to more than {t} euros?",
]
def make_case(i):
prices = [round(rng.uniform(5, 80), 2) for _ in range(rng.randint(2, 4))]
total = sum(prices)
gap = rng.choice([1, 2, 5, 15]) # near and far from the threshold
t = int(total) - gap if i % 2 == 0 else int(total) + gap
state = {"items": [{"name": f"item {k + 1}", "price_eur": p} for k, p in enumerate(prices)]}
question = rng.choice(TEMPLATES).format(t=t)
request = {"model": "jev-latest", "state": state,
"questions": {"over": {"type": "noul", "instructions": question}}}
return request, {"over": "yes" if total > t else "no"}
out = pathlib.Path("cases"); out.mkdir(exist_ok=True)
for i in range(500):
request, expect = make_case(i)
(out / f"{i:04d}.request.json").write_text(json.dumps(request))
(out / f"{i:04d}.expect.json").write_text(json.dumps(expect))Three details carry most of the value. The label is computed from total > t, not from which
branch was taken, so a bug in the branch logic cannot produce a wrong label. The gap puts some
cases within a euro of the threshold and some far away, so you can see whether errors cluster at
the boundary. And alternating the branch keeps the yes and no answers balanced.
- Balance by construction. Decide the answer first, then build the text that produces it, and check it with the computed label. A generator that picks values at random and lets the answer fall where it may will produce whatever ratio the ranges imply.
- Boundaries on purpose. Include cases exactly at the threshold ("over 100" with a total of 100.00) and decide in the question whether "over" includes it. These are the cases where question wording and label logic most often disagree.
- Distractors. Add values that should be ignored: a shipping fee the question does not mention, a date that is not the delivery date, a second customer's order. Without them, the only number in the text is always the relevant one, and the test is easier than real life.
- Missing information. Leave the relevant field out of some texts and ask whether the text says it. These "not stated" cases are cheap to generate and catch confident answers to unanswerable questions; the question pattern is on ask whether the text says it at all.
A single template tests one phrasing, and models can be sensitive to wording. Write three or more templates per question, tag each case with its template, and report accuracy per template as well as per kind. A large gap between templates is a finding about wording, not about the model's ability; how to use it is on why wording changes an LLM's answer.
- Fix the seed and store it with the results.
- Version the generator with your code. A changed template is a changed test.
- Record the generator version, the seed, the model file and its hash next to every reported score, so a number from last month can be compared with one from today.
- Keep generated and hand-labelled results in separate reports. Averaging them hides which one moved.
Being straight about the limits:
- Clean text. Templates state numbers plainly; real emails bury them in prose. In our own two test sets, number-against-threshold questions scored 0.85 on the generated set and 0.654 on the hand-written one, most likely for this reason. Arithmetic (0.56 and 0.584) and dates (0.61 and 0.598) agreed closely. The comparison is on small LLMs and arithmetic in yes/no questions.
- A style of its own. Every text from a template shares a structure, and a system tuned against that structure can do better on the generated set than on real data. Use it to test, and if you ever tune on something, tune on anything but this.
- Only what you thought of. A generator covers the variations its author imagined. Real inputs contain the ones nobody did, which is why a hand-labelled sample of real cases stays necessary.
- Correct labels for the question you wrote. If the template says "within 30 days" and your business means "within 30 calendar days of delivery, excluding the delivery day", the code will faithfully compute the wrong thing. Review the label logic like any business rule.
How do I create test data for an LLM without labelling? Generate texts from templates and compute each answer in code from the values you inserted.
Are generated test questions as good as real ones? They are more reliable in their labels and less realistic in their text. Use both: generated for size and exact answers, real for coverage.
Which kinds of question can be generated? Anything whose answer follows from structured values: thresholds, sums, dates, explicit rules, and whether a field is present.
How do I avoid an easy generated test? Add distractor values, missing fields, cases at the exact threshold, and several phrasings per question.
Did generated and hand-written tests agree in your measurement? On arithmetic and dates, within about two points. On thresholds, no: 0.85 generated against 0.654 hand-written.
See also: building a yes/no test set for your own data, accuracy by kind of question and benchmark contamination and truly held-out tests.
- Per-kind accuracy on our generated 1,000-question test set (answers computed by code, over orders, leave, servers, loans, courses, shipments, rentals, prescriptions and bookings) and on our 999-question hand-written set: our measurements of the released jevos.
- The generator is an illustrative sketch written for this page.
From the notes of jev, where the generated test and the hand-written one agreed on arithmetic and disagreed on thresholds, and both results were kept.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building