-
Notifications
You must be signed in to change notification settings - Fork 130
llm regression tests in ci
An LLM regression test in CI is a fixed set of inputs, your system's outputs for them, and a
list of yes/no checks on each output, asserted as thresholds on P(yes). With jevos, jev decide answers one request file without a server, so a CI step can judge every output on a CPU
runner and fail the build when a check that used to pass drops below its bar. The judge adds no
per-token cost and no data leaves the runner.
What makes these tests usable is the margin. A check that asserts P(yes) above 0.5 will flip on and off for outputs that sit near 0.5, and a test suite that flips is a test suite people learn to ignore. The fix is a pass band, a fail band and a warning band in between.
This page is the shape of a check, the files and the loop, assertions with margins, the causes of flaky results, how to pin the judge so its answers are comparable over time, and what to keep out of CI.
Two things change in an LLM application: your prompts, retrieval and code, and the model behind them. A regression test catches either one making outputs worse on properties you care about:
- "Does the reply ask the customer for their order number?"
- "Does the reply promise a refund?" (which it must not, for this input)
- "Is the tone of the reply polite?"
- "Does the summary mention the cancellation date?"
Each is a property you can name and read in the output, which is the kind of criterion a small judge handles well. How to write them is on rubric design for an LLM judge. A check such as "is the total in the reply correct?" is arithmetic, and belongs in an ordinary unit test that compares numbers.
A CI job has two stages. First your application produces outputs for a fixed list of inputs and writes one request file per case. Then the judge answers each file.
{
"model": "jev-latest",
"state": {
"customer_message": "My parcel never arrived and the tracking has not moved in a week.",
"reply": "Sorry about that. Could you send me your order number so I can check with the courier?"
},
"questions": {
"asks_order_number": {"type": "noul", "instructions": "Does the reply ask the customer for their order number?"},
"promises_refund": {"type": "noul", "instructions": "Does the reply promise the customer a refund?"},
"polite": {"type": "noul", "instructions": "Is the tone of the reply polite?"}
}
}Keep the expectations out of the request, in a sidecar file per case, so the request stays the
documented body of POST /v1/systemone:
{"asks_order_number": "yes", "promises_refund": "no", "polite": "yes"}mkdir -p answers
for f in cases/*.request.json; do
name=$(basename "$f" .request.json)
jev/jev decide "$f" --output "answers/$name.json"
done--output writes to a new file and never overwrites an existing one, so start each job with an
empty answers directory. The answer file has the same shape as a server response, with one
probability per question under answers.
Each run of jev decide starts from the model file. For a suite of hundreds of cases it can be
simpler to start jev serve once in the job, wait until GET /health reports ready, and post the
files to it. We have not measured the start-up cost of jev decide, so measure it on your runner
before choosing. Runner setup, downloading the model and caching it, is on
running LLM yes/no checks in GitHub Actions.
import json, pathlib, sys
PASS, FAIL = 0.7, 0.3 # tune on your own labelled cases
failed, warned = [], []
for exp in pathlib.Path("cases").glob("*.expect.json"):
name = exp.name.removesuffix(".expect.json")
got = json.loads(pathlib.Path(f"answers/{name}.json").read_text())["answers"]
for q, want in json.loads(exp.read_text()).items():
p = got[q]["noul"]
ok = p >= PASS if want == "yes" else p <= FAIL
bad = p <= FAIL if want == "yes" else p >= PASS
if bad:
failed.append(f"{name}.{q} = {p:.2f}")
elif not ok:
warned.append(f"{name}.{q} = {p:.2f}")
print("\n".join(["FAIL " + x for x in failed] + ["WARN " + x for x in warned]))
sys.exit(1 if failed else 0)Three outcomes per check: clearly right (pass), clearly wrong (fail the build), and in between (warn, and show it in the job summary). The warning band is where a person should look, and it is also where the judge itself is least reliable. Choosing the two numbers from labelled cases, rather than guessing them, is the subject of how to choose a threshold for P(yes).
Make the bands asymmetric where the error costs are. jevos leans toward yes on questions it cannot work out, so an expected-yes check deserves a higher pass bar than an expected-no check deserves a low one.
Most flakiness in these suites does not come from the judge:
- The system under test samples. If your generator runs with a non-zero temperature, the same input gives different outputs, and a check can legitimately pass on one and fail on another. Either fix the generator's sampling for the test run, or generate several outputs per case and assert on the pass rate.
- Outputs near a threshold. Handled by the warning band.
- Checks that are really two checks. "Is the reply polite and does it ask for the order number?" fails for two reasons and passes for one. Split it.
- A changed judge. If the model file or the runtime changes, scores move. See the next section.
When a check flips, look at the output before touching the threshold. A threshold nudged until the build is green is a test that no longer tests anything.
A regression suite compares today's outputs with yesterday's, and that only works if the judge is
the same. Pin the model file by name and check it against the release's SHA256SUMS.txt in the
job. If you use jev serve, GET /health reports the model's fingerprint, which you can print
into the job log. Upgrade the judge in its own commit, re-run the
suite, and review the scores that moved, the same way you would review a dependency upgrade.
- The final evaluation of a release. CI checks catch regressions; they do not tell you how good the system is. That needs a proper test set, described on building a yes/no test set.
- Checks the judge is weak on. Numbers, dates and sums go into ordinary code assertions.
- Non-English outputs. jevos reads English only.
- Latency benchmarks on shared runners. CI machines are shared and noisy; a judge timing from one run says little.
How do I test LLM outputs in CI? Write yes/no checks for the properties you care about, run a judge on each output, and assert on the probabilities with a pass band, a fail band and a warning band between them.
Do I need a GPU in CI? No. jevos runs on the CPU with about 1 GB of memory.
Why do my LLM tests pass and fail at random? Usually because the system under test samples, or because outputs sit near a single threshold. Fix the sampling and add a warning band.
Should the judge run as a server or per file? jev decide needs no server; for large suites,
one jev serve per job may be simpler. Measure both on your runner.
How do I know the judge did not change? Pin the model file, verify it against
SHA256SUMS.txt, and log the fingerprint that /health reports.
See also: batch decisions from files with jev decide, LLM as a judge on a CPU and LLM judge bias and how to control it.
-
jev decide,--output,/healthandSHA256SUMS.txt: the jev README and the jevos release. - The lean toward yes: our 999-question test set,
jevos-q4_k_m. - The shell and Python snippets are sketches written for this page, not tested scripts.
From the notes of jev, whose --output flag never
overwrites an existing file, which is the behaviour you want from a test result.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building