-
Notifications
You must be signed in to change notification settings - Fork 132
llama cpp cpu only
To run llama.cpp on the CPU only, either install a CPU build or tell it to offload nothing
(--device none or -ngl 0 on llama.cpp's tools), then set the thread
count to roughly your number of physical cores and measure from there. llama.cpp's own docs
recommend physical rather than logical cores for --threads. jev runs on the CPU only, so it has
nothing to switch off; its --threads defaults to all logical CPUs, and fewer is better if other
heavy apps are running.
The part that is not obvious is that "use every core" is a starting point, not an answer. On processors that mix fast and slow cores, and on machines doing other work, the best thread count is often lower than the maximum. It depends on the machine, so it has to be measured.
This page is the switches that keep llama.cpp off the GPU, how the thread options work, why hybrid cores make the choice harder, what memory looks like on the CPU, and when a CPU is not enough.
There are two independent levers, and it helps to know both.
Which runtime you install. llama.cpp releases have CPU-only archives for each platform. A CPU build has no GPU backend to pick, so nothing can be offloaded by accident.
Which devices you ask for at run time. A GPU build can still run entirely on the CPU:
-
llama-serverand the other tools take--device none, documented as "none = don't offload", or-ngl 0to keep no layers in VRAM.--list-devicesshows what is available. - jev has no device option: it runs its 8-bit OpenVINO model on the CPU only and does not load llama.cpp's runtime, so there is no GPU to keep it off.
./jev serve --threads 16llama.cpp's tools have two thread settings:
| Option | What it controls | Default |
|---|---|---|
-t, --threads
|
threads used during generation | -1 |
-tb, --threads-batch
|
threads used for batch and prompt processing | same as --threads
|
The docs add that "in some systems, it is beneficial to use a higher number of threads during
batch processing than during generation", and recommend physical cores for --threads.
For a decision model the second one is the one that matters. jevos generates no tokens, so all
of its time is prompt processing, the phase explained on
prefill vs decode. jev has a single --threads setting, so
there is only one knob to tune.
Many recent laptop and desktop chips mix core types. The reference laptop for the jevos
numbers, an Intel Core Ultra 7 255H, has 16 cores in three kinds, 6 performance cores, 8
efficient cores and 2 low-power efficient cores, and 16 threads in total. The published
measurements used --threads 16.
The theory of why more threads can be slower, which we have not measured on this page:
- Each step waits for the slowest thread. Work on a matrix is split across threads, and the step finishes when the last share is done. A share given to a slower core, or to a core that another program is using, holds up the rest.
- Memory bandwidth is shared. Past some point, more cores are waiting on the same memory rather than computing. Why memory traffic dominates on a CPU is covered on what makes a local LLM fast on a CPU.
- Other work competes. A browser, a build or a video call takes cores away without telling llama.cpp, which is why the README says "fewer if other heavy apps are running".
llama.cpp's tools also offer placement controls for this: --cpu-mask for an affinity mask,
--cpu-strict for strict placement, --prio for priority and --poll for how threads wait
for work. jev does not expose these; it sets the thread count only.
The practical answer: try a few values (for example the number of performance cores, then more), and compare medians after warm-up, one run at a time. The method is on measuring LLM latency: median, p90 and warm-up.
On the CPU the model lives in ordinary RAM, and llama.cpp's default is to memory-map the file:
the load mode auto maps the model "unless the device does not support it". Mapping lets the
operating system page weights in from the file instead of copying them up front. The
alternative mlock pins the model in RAM so it cannot be swapped out; the docs note it "can
improve performance but trades away some of the advantages of memory-mapping by requiring more
RAM to run and potentially slowing down load times."
For the jevos GGUF files the numbers are small: the q4_k_m file is 619 MB and memory use grows by about 1.2 GB with it loaded in llama.cpp, with a context of up to 8,192 tokens. That fits alongside normal work on most machines, which is what makes CPU-only practical for it.
On the CPU, jev on the reference laptop answers a short request (about 30 tokens) in 26 ms and a long one (about 190 tokens) in 112 ms, each read from scratch. Cost grows with the length of the text, so the CPU is comfortable for messages, tickets and records, and slower for long documents.
Being straight about the limit: a CPU is the wrong tool when you process long documents at volume or need high throughput from one box. Then a GPU build of llama.cpp running the jevos GGUF files is the right kind of tool (jev itself has no GPU path), and the trade-off is on CPU or GPU for a small LLM.
How do I force llama.cpp to use only the CPU? Use a CPU build, or pass --device none or
-ngl 0 to llama.cpp's tools. jev runs on the CPU only, with nothing to set.
How many threads should llama.cpp use? Start at the number of physical cores, as llama.cpp's docs recommend, then measure lower values. Fewer if other heavy programs are running.
What is the difference between --threads and --threads-batch? The first is for generation, the second for batch and prompt processing. A decision model only does the second.
Does llama.cpp use efficient cores? It runs threads wherever the operating system places them unless you set an affinity mask. Whether that helps on your chip is a measurement.
Why is my CPU inference slower with more threads? Usually shared memory bandwidth, slower cores holding up each step, or other programs using the same cores.
See also: llama.cpp on Windows without compiling, run an LLM locally without a GPU and an LLM on a laptop.
- Our own measurements and facts: latency, tokens, memory and context of jevos on the reference
laptop with 16 threads, and the
--threadsguidance and default, from the jev README. -
llama-server README,
fetched 2026-09-29:
--threads,--threads-batch,--device,-ngl,--list-devices,--cpu-mask,--cpu-strict,--prio,--poll. -
llama.cpp completion README,
fetched 2026-09-29: physical-core recommendation, batch threads note, load modes and
mlock. - Intel Core Ultra 7 255H specifications, fetched 2026-09-29: core types and thread count.
From the notes of jev, whose published timings all come from one hybrid-core laptop with its 16 threads in use.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building