-
Notifications
You must be signed in to change notification settings - Fork 131
cpu or gpu for a small llm
For a small model reading short inputs one request at a time, a modern CPU is usually enough, and it is the hardware you already have. A GPU pays off when there is a lot of parallel arithmetic to do: many requests batched together, long documents, or long generated answers. jev is built and measured for the first case: it runs on the CPU only, measured on a laptop, and has no GPU path.
The question is less "which is faster" than "which is idle". A GPU that finished a 50 ms job in 10 ms (illustrative numbers, not a measurement) would save 40 ms per request, which matters if you have thousands of requests per second and barely matters if you have ten a minute. What decides it is your traffic, not the hardware's peak.
This page is three questions that decide it, what each processor is good at, the costs that are not about speed, how jev picks a device, and when to switch.
- How many decisions per second at peak? A handful per second with no queue is a CPU job. Sustained high concurrency is where batching on a GPU earns its cost.
- How long are the inputs? Tens to a few hundred tokens read quickly on a CPU; on our laptop jevos took 26 ms at about 30 tokens and 112 ms at about 190. Thousands of tokens per request multiply that.
- Does the model generate? A decision model with zero output tokens avoids the long sequential phase entirely. A chat model writing paragraphs spends most of its time there.
If the answers are "few", "short" and "no", stay on the CPU and spend the effort on the input instead. The factors that set CPU speed are on what makes a local LLM fast on a CPU.
Language model inference has two phases. The Splitwise paper describes the prompt phase as compute-intensive, many tokens processed in parallel, and the token generation phase as bound by memory bandwidth and capacity. GPUs are built for the parallel arithmetic of the first phase; how much they help the second depends on the memory they read from.
The same paper notes that batching the token phase "yields high throughput", and the vLLM paper opens with "high throughput serving of large language models (LLMs) requires batching sufficiently many requests at a time". Both statements are about serving many users. A GPU's advantage grows with the amount of work it can do at once, which is why it shines under load and matters less for one short request.
What a GPU does not remove: the time your application spends building the request, the HTTP round trip, and any work your code does with the answer. For a request that already takes tens of milliseconds, those parts are a larger share than they look.
- Availability. Every server and laptop has a CPU. GPUs in a data centre are a separate budget line; on a laptop, the integrated one may or may not be supported by your runtime.
- Memory. On a CPU the loaded model lives in system memory, which most machines have spare; on a GPU it has to fit in the card's memory next to whatever else runs there.
- Deployment. A CPU-only service is one binary and one model folder on any machine; a GPU service adds drivers and a matching runtime build. llama.cpp publishes builds for many backends (CUDA, HIP, Metal, Vulkan, SYCL and others, per its README), which helps, but it is still one more thing to match.
- Contention. A shared GPU is shared latency. So is a shared CPU; see the note on benchmarking on measuring LLM latency.
For many teams the deciding argument is that the CPU servers already exist; that angle is on on-premise LLM for business decisions.
It does not pick one: jev runs on the CPU only. It runs jevos-v2 with 8-bit (INT8) weights through OpenVINO, on any x86-64 CPU with AVX2 and on Apple silicon, and it is fastest on CPUs with AVX-VNNI or AVX-512 VNNI. There is no device option and no GPU build.
Every jevos number we publish was taken on an Intel Core Ultra 7 255H with 16 threads, the configuration of the README's benchmark. The release also ships the model as GGUF files for llama.cpp and other tools; if you run those on a GPU, measure it yourself: we make no claim about the speed you will see, and the CPU-only guide for llama.cpp covers the flags that matter when you stay on the CPU.
- Throughput beyond one machine's CPU. When requests queue and the p90 climbs past your budget, the choices are more CPU processes, more machines, or a GPU with batching. The trade-off is on throughput vs latency for a decision server.
- Long documents at interactive speed. Thousands of tokens per request is prompt-heavy work, the phase GPUs accelerate most.
- A larger model. If a small model is not accurate enough for your questions, the model you move to may not be practical on a CPU at all. llama.cpp's "CPU+GPU hybrid inference" exists for models "larger than the total VRAM capacity", in its own words, which says something about the sizes involved.
Being straight about the limit: a small model on a CPU is a good fit for a narrow job. It is not a way to avoid GPUs for everything.
Do I need a GPU to run an LLM? Not for a small model on short inputs. jevos is measured on a laptop CPU with no GPU in use. See run an LLM locally without a GPU.
Is a GPU always faster? Per unit of work it usually is. For one short request the difference can be small next to everything else in the request, and the GPU has its own costs.
Can jevos use a GPU? jev runs on the CPU only. The GGUF files in the release can be run elsewhere, for example in llama.cpp; we have only published CPU measurements.
What about an integrated GPU? jev does not use it. If you run the GGUF files on one, whether it beats the CPU cores on your laptop is something to measure, not assume.
When is the CPU the wrong choice? High sustained concurrency, very long inputs, generation, or a model too large to run well in system memory.
See also: the fastest AI model for yes/no decisions, edge AI decisions on a CPU and an LLM on a laptop.
- jevos latency, the reference machine and the CPU requirements: our measurements and the jev README.
- The two inference phases and batching the token phase: Patel et al., Splitwise, fetched 2026-09-29.
- Batching for throughput: Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, fetched 2026-09-29.
- Backends and hybrid inference: llama.cpp README, fetched 2026-09-29.
From the notes of jev, which runs on the CPU only, the one setting all of its numbers were measured on.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building