-
Notifications
You must be signed in to change notification settings - Fork 131
llama cpp vs ollama for classification
For a classification service, llama.cpp's llama-server gives you more control over how the
model runs, and Ollama gives you easier model management; neither returns a class
probability out of the box, so with both you generate one token and read its probability.
Ollama's README credits llama.cpp as its supported backend, so the difference is less the
engine than the layer on top: flags and a pinned release on one side, ollama pull, a model
store and automatic loading and unloading on the other.
Conflict of interest: we build jevos and jev serve, a dedicated decision server for
exactly this job. We build neither llama.cpp nor Ollama, and both are good at what they are for.
The point that decides most choices is not speed but the contract. A classifier needs the same prompt, the same model file and the same parsing on every call, and a number at the end. Both servers are designed around chat and text generation, so that contract is something you build on top.
This page is what each project is, what a classification endpoint needs, how to get a probability from each, the operational differences, and when neither is the simplest answer.
llama.cpp describes itself as "LLM inference in C/C++". Its llama-server is a "fast,
lightweight, pure C/C++ HTTP server" with OpenAI-compatible chat completions, responses and
embeddings routes, a native /completion route, /tokenize, /reranking, /health,
continuous batching, parallel slots and schema-constrained JSON output. You choose the model
file, the context size (-c), the threads (-t), the devices (--device) and the number of
slots (-np) on the command line.
Ollama is a model runner with a CLI (pull, run, create, serve, ps, list, rm)
and a REST API on port 11434, with /api/generate and /api/chat. It manages a local model
store, imports GGUF or safetensors through a Modelfile, and loads models on demand. Its FAQ
states the defaults that matter for a service: it binds 127.0.0.1:11434, keeps a model in memory
for 5 minutes after use, processes 1 request per model at a time
(OLLAMA_NUM_PARALLEL), and uses a 4096-token context unless told otherwise.
Whatever the server, a yes/no or label decision needs five things:
- A fixed prompt. The text and the question are placed in the same template every time.
- A number, not a word. You want P(yes), so you can set a threshold; parsing "Yes.", "yes" or "Yes, because..." is the fragile part.
- Cheap extra questions. Several questions about one text should not each pay to read the text again.
- A pinned model. The file hash and the runtime version should be known for every answer.
- Predictable latency. No cold load in the middle of traffic.
Neither server offers items 2 and 3 as a ready-made classification feature; both give you the pieces.
With llama-server, the native /completion route takes n_probs, which returns "the
probabilities of top N tokens for each generated token", and post_sampling_probs turns them
into probabilities between 0 and 1 after the sampling chain. You ask for one token, read the
probability of the "yes" token and of the "no" token, and normalize.
With Ollama, /api/generate accepts logprobs ("whether to return log probabilities of
the output tokens") and top_logprobs. Same approach: one output token, read the candidates.
The pieces you then write yourself, on either server, are the same:
- the prompt template, and
rawmode or the model's chat template applied consistently; - the list of tokens that mean yes and no for that tokenizer (with and without a leading space, upper and lower case), since a probability split across variants undercounts both;
- a fallback when neither yes nor no is among the top candidates;
- one call per question, unless you manage prompt caching yourself.
That is the gap jev serve fills; the detail of what you would build on llama-server is on
jev serve vs llama.cpp server. The underlying reason one
forward pass is enough, and generating is not needed, is on
why one forward pass beats generating an answer.
| llama-server | Ollama | |
|---|---|---|
| Model source | a GGUF path, or -hf from the Hub |
ollama pull or a Modelfile import |
| Context | from the model unless -c is set |
4096 tokens by default |
| Idle model | not documented as unloaded | unloaded after 5 minutes by default |
| Concurrency | parallel slots, continuous batching | 1 parallel request per model by default |
| Probabilities |
n_probs on /completion
|
logprobs, top_logprobs on /api/generate
|
Two rows matter most for classification. The context default: a long document plus
instructions can exceed 4096 tokens, so on Ollama set the context explicitly. The idle
unload: a classifier that runs every few minutes will pay a model load on some requests unless
keep_alive is raised. Both are one setting; both are easy to miss.
-
Ollama fits if you already run it for chat, want several models on one machine, or value
pulland automatic memory management over flags. A handful of classifications throughlogprobsis reasonable there. - llama-server fits if you want one pinned model file and one pinned release per service, explicit threads and devices, and parallel slots under load.
-
A dedicated decision server fits if the whole job is yes/no answers. jevos returns P(yes)
per question with
output_tokensalways 0, reads the shared text once for several questions (three questions take about 66 ms against 49 ms for one on the reference laptop), and reports the served model in/health. It is English only and answers yes/no, multiple-choice and (early) score questions only; a comparison at the product level is on jevos vs Ollama for yes/no decisions.
Does Ollama use llama.cpp? Ollama's README lists "llama.cpp project founded by Georgi Gerganov" under supported backends.
Can llama-server return class probabilities? Not as classes. /completion with n_probs
returns the probabilities of the top tokens for each generated token; you map tokens to labels.
Can Ollama return logprobs? Yes. /api/generate takes logprobs and top_logprobs.
Which is faster for classification? There is no general answer. It depends on the file, the context, the threads and whether the model was already loaded. Measure both on your hardware with the same GGUF and the same prompt.
Why not just ask the model to answer "yes" or "no" in text? You can, but then you parse text, lose the confidence, and pay for decoding.
See also: structured output vs a probability, using llama.cpp prebuilt binaries and zero-shot text classification with yes/no questions.
- Our own facts: jevos outputs one probability per question with zero output tokens; the 66 ms
and 49 ms timings on the reference laptop;
/healthcontents. From the jev README and source code. -
llama.cpp repository and
llama-server README,
fetched 2026-09-29: features, endpoints,
n_probs,post_sampling_probs, command-line options. -
Ollama README, Ollama FAQ,
Ollama generate API and
Ollama import guide, fetched 2026-09-29: CLI, port, defaults
for keep-alive, parallelism and context,
logprobs. -
Hugging Face docs: GGUF with llama.cpp,
fetched 2026-09-29: the
-hfoption.
From the notes of jev, a decision server whose model also ships as GGUF files that either of them can load.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building