-
Notifications
You must be signed in to change notification settings - Fork 131
small language models explained
A small language model is one with few enough parameters, roughly a billion rather than tens or hundreds of billions, to run on ordinary hardware such as a laptop CPU. What it gives up is breadth: it reads well and reasons poorly. On 999 yes/no questions written after it was finished, the 1B model jevos was right 0.954 of the time on facts stated in the text and 0.938 on tone, but 0.584 on questions that needed arithmetic and 0.598 on dates. Small models suit narrow tasks where the job is to read, and where computing can be done in code.
"Small" is not a technical category with a fixed boundary. It is a practical one: small enough that the hardware stops being the project, and that one person can measure the model's behaviour on their own data in an afternoon.
This page is what "small" means, what a small model is good and bad at with measured numbers, why a narrow task changes the picture, how small compares with large on the same job, and why small-model benchmarks deserve suspicion.
There is no official threshold. In practice people call a model small when it runs without specialised hardware: a few hundred million to a few billion parameters, quantized to a few bits per weight, in a file of hundreds of megabytes to a few gigabytes. jevos is at the small end of that range: about 1 billion parameters, which jev runs with 8-bit (INT8) weights, and its 4-bit GGUF file is 619 MB.
Size sets the cost of every token processed. It also sets how much the model can know and how many steps of reasoning it can hold together, which is where the trade-off lives.
Reading. Questions whose answer is in the text, in some form, are where a small model is strongest. From the 999-question set, by kind of question:
- stated fact: 0.954 (108 questions)
- tone: 0.938 (32)
- paraphrase, do two phrasings mean the same: 0.893 (84)
- the writer's intent: 0.859 (71)
- negation, the text says something is not so: 0.858 (106)
- not stated, the text does not say it at all: 0.847 (98)
These are the questions most applications actually ask: is this a billing problem, is the customer upset, does the email ask for a meeting, does the log line mention a customer-facing service. The texts in the set were emails, tickets, logs, reviews and forms of 40 to 150 words, and none of the questions was used to tune anything.
Computing. The same set, the other end:
- applying a written rule: 0.721 (104)
- a number against a threshold: 0.654 (78)
- dates and durations: 0.598 (97)
- arithmetic: 0.584 (221)
The failure has a direction. When the model cannot work out the answer it leans toward yes: 152 of its mistakes were a yes that should have been no, against 91 the other way, and on arithmetic questions whose answer is no the mean P(yes) was 0.59. A small model is not a random guesser on these questions; it is a biased one. The measurement is on why a small LLM says yes when the answer is no, and the practical fix, extracting the numbers and comparing them in code, is on small LLMs and arithmetic in yes/no questions.
Because a narrow task removes the parts of the job a small model does worst, and makes the rest measurable.
-
No format to get wrong. A model that returns one probability per question has no JSON to
break, no "Yes, because..." to parse and no answer in the wrong language.
output_tokensis always 0. - No long chain of reasoning. A yes/no question about one condition is one step. Splitting a compound decision into several such questions, and combining the answers in code, keeps each step inside what the model reads well.
- A task you can test completely. With one kind of output, you can build a set of your own labelled cases, split it by kind of question, and know where the model is and is not reliable before it touches production. Accuracy by kind of question explains why the split matters more than the overall number.
A general assistant has none of these properties. That is why a model that is weak as a chatbot can be useful as a component.
On 2,000 yes/no questions about three business policies that none of the models was tuned on, with answers computed by code, the hosted Jev was right 0.927 of the time and jevos 0.810. The gap was largest on additive point scores, where several signals are summed and compared with a cut-off: a computation again.
The other side of the trade is speed and place. On the same two requests, jevos took 26 and 112 ms on a laptop CPU; the hosted API took 344 and 345 ms from Europe, network included. The full comparison is on jevos vs Jev vs Laya. A small model is the cheaper, faster, local first reader; a large one is the better judge of hard cases. The choice between them is laid out on when a small model is enough.
A benchmark built like the data a model was developed on measures familiarity, not generalisation. On a held-out split built the usual way, jevos scored 0.855 overall, with a calibration error of 0.009 on natural yes/no questions. On the 999 questions written from scratch afterwards, it scored 0.757. The gap is the subject of our held-out benchmark said 0.855. The lesson applies to any small model: trust a test built from your own cases over any published number, including ours.
What is a small language model? A language model small enough, roughly a billion parameters, to run on ordinary hardware without a GPU. There is no fixed cut-off.
Are small language models accurate? At reading, fairly: 0.954 on stated facts in our test. At computing, poorly: 0.584 on arithmetic. The kind of question matters more than the model.
What are small models used for? Narrow tasks such as classification, routing, filtering and yes/no checks, where the answer is in the text and the output is simple.
Can a small model replace a large one? For many reading decisions, yes. For rules, sums, dates, other languages or open-ended answers, keep the large model or move the logic into code.
Why does a small model say yes when it does not know? In our measurement it leans toward yes on questions it cannot compute. Keep computation out of the question.
See also: run an LLM locally without a GPU, how to write yes/no questions an LLM answers well and LLM as a judge on a CPU.
- All accuracy, error-direction and calibration figures are our own measurements: the
999-question set on
jevos-q4_k_m, the held-out split onjevos-q8_0, and the 2,000-question policy comparison. - Latency on the two requests: our measurements, reported in the jev README.
From the notes of jev, a small model we describe by where it fails as much as by where it works.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building