-
Notifications
You must be signed in to change notification settings - Fork 132
model cascade small model first
A model cascade asks the small model first, keeps its answer when the probability is clearly high or clearly low, and sends only the uncertain middle to a larger model. With a local yes/no model the first stage costs 25 to 110 ms on a laptop CPU and nothing per token, so the large model is paid only for the cases that need it. Because jevos speaks the same wire format as TypeSafe's hosted Jev, the escalation can be the same request body sent to a different base URL.
A cascade differs from a router in one way: a router decides where a request goes before any model answers, and a cascade lets the cheap model try and uses its own confidence to decide whether to stop. That makes the cascade only as good as the small model's probabilities, and this page is honest about where they are and are not a good signal.
This page is the shape of a cascade, the escalation step, the cost and latency arithmetic, the case where confidence misleads, and how to set the band from your own labelled cases.
For each yes/no question, the local answer lands in one of three bands:
| P(yes) from the small model | What the cascade does |
|---|---|
| at or above the high bar | act on yes, done |
| at or below the low bar | act on no, done |
| in between | ask the large model, act on its answer |
The bars are yours to set, per question. A symmetric band such as 0.2 to 0.8 is a place to start, not a recommendation. Where a wrong yes is the expensive mistake, move the high bar up; the reasoning is on how to choose a threshold for P(yes).
The idea is not new. FrugalGPT, a 2023 paper by Chen, Zaharia and Zou, describes LLM cascades that learn which combination of models to query for each input, and reports matching the performance of the strongest model it tested with up to 98% cost reduction on its benchmarks. Those are their numbers on their tasks, not a prediction for yours.
jevos accepts TypeSafe Jev's request format: model, state, and named noul, choice or
score questions. Code written for Jev's SDK works unchanged for all three. So the second stage does not need a
second integration:
import requests
LOCAL = "http://127.0.0.1:8017/v1/systemone"
def decide(body, hosted_url, low=0.2, high=0.8):
local = requests.post(LOCAL, json=body, timeout=2).json()["answers"]
unsure = {k: q for k, q in body["questions"].items()
if low < local[k]["noul"] < high}
if not unsure:
return {k: v["noul"] for k, v in local.items()}, "local"
second = requests.post(hosted_url, json={**body, "questions": unsure}, timeout=10).json()["answers"]
merged = {k: v["noul"] for k, v in local.items()}
merged.update({k: v["noul"] for k, v in second.items()})
return merged, "escalated"Only the uncertain questions are sent on. Authentication for the hosted service is left out of
the sketch. Two limits apply: the sketch reads only noul answers (jevos also answers choice
and score questions, each with a confidence to cascade on), and the large model does not have to be Jev. Any model you trust more can
be the second stage; the shared wire format only makes Jev the one with no extra code.
Let f be the fraction of questions that land in the middle band. Then, per request:
- expected latency is about the local time, plus f times the hosted time;
- expected paid calls are f times what you paid before, since the confident ends cost nothing per token.
With our measured numbers for a short request, 26 ms locally and 344 ms for the hosted Jev from Europe with the network included, the arithmetic looks like this. The escalation fractions are illustrative, not measurements:
| Escalated fraction f (illustrative) | Mean latency | Paid calls, relative to all-hosted |
|---|---|---|
| 0.1 | 26 + 34 = about 60 ms | 0.1 |
| 0.3 | 26 + 103 = about 129 ms | 0.3 |
| 0.6 | 26 + 206 = about 232 ms | 0.6 |
Past a certain f the cascade is slower than asking the large model directly, because every escalated case pays both. The tail matters too: the slowest requests are the escalated ones, so the p90 of a cascade can be worse than its mean suggests. Measure f on your own traffic before deciding.
A cascade assumes that a wrong answer comes with a middling probability. That holds where the model is calibrated. On the natural yes/no questions of our held-out split, the calibration error was 0.009, which means confident answers there are right about as often as they claim.
It does not hold everywhere. On 999 questions written after training, the model made 152 mistakes by saying yes when the answer was no, against 91 the other way, and the lean sat on arithmetic and dates. On arithmetic questions whose answer was no, the mean P(yes) was 0.59. A confident wrong answer does not land in the middle band, so the cascade never escalates it. The measurement is on why a small LLM says yes when the answer is no.
Two fixes, both in design rather than in thresholds:
- Route by kind of question as well as by confidence. A question that needs a computation should not go to the small model at all. Compute it in code, or send it straight to the large model.
- Keep the yes bar higher than the no bar. On our sets the small model's no was the more reliable answer.
- Collect a hundred or more real cases per question, labelled by hand.
- Run the small model on all of them and record P(yes).
- For each candidate band, count how many cases fall outside it (answered locally) and how many of those are wrong.
- Pick the narrowest band whose local error rate you can accept, and read f off the same table.
This is the same exercise as sizing a human review band, with a large model in the place of the person. The two can be stacked: small model, then large model, then a person for what is still unclear.
- The large model is needed for most cases. If f is high, you pay two latencies for little saving. Ask the large model directly.
- The task is not yes/no. Extraction, summaries and replies are generated text; the local stage has nothing to offer them.
- The data may not leave the machine. Then the second stage has to be local too, or a person.
What is a model cascade? A chain of models where a cheap one answers first and a more expensive one is asked only when the first is unsure.
How is it different from a router? A router picks the model before anyone answers. A cascade uses the first model's own confidence to decide whether to go on.
Does it make every request faster? No. Confident cases are faster; escalated ones pay both latencies.
Can I escalate from jevos to Jev without new code? For yes/no questions, the request body is the same, so the escalation is a second POST to another base URL.
What goes wrong with cascades? Confidently wrong answers are never escalated. On our tests those cluster on arithmetic and dates.
See also: an LLM router with yes/no questions, local vs hosted LLM decisions and reducing LLM cost with local yes/no decisions.
- Our measurements: 26 ms local and 344 ms hosted on the same short request, the 0.009 calibration error on the held-out split, and the error direction and mean P(yes) on the 999-question set. Latency figures are in the jev README.
- Wire format compatibility and the
scoreanswers: the jev README. - FrugalGPT: Chen, Zaharia and Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, arXiv 2305.05176, fetched 2026-09-29.
From the notes of jev, which answers the easy end of a cascade. The hard part of a cascade is the case it never escalates, so we measured that first.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building