-
Notifications
You must be signed in to change notification settings - Fork 131
jev serve vs llama cpp server
llama.cpp's llama-server is a general server for completions, chat, embeddings and reranking;
jev serve runs one yes/no model on the CPU behind one decision endpoint, POST /v1/systemone, that returns P(yes) per question. Both are local, and
llama-server gives you every building block a decision service needs. What it does not give you is
the service: the prompt, the reading of a yes/no probability from token alternatives, several
questions sharing one text, and a stable request and answer format. If you want those built,
jev serve has them; if you want full control or a different model, llama-server is the
foundation to build on.
Conflict of interest, in one line: we build jev; the llama-server facts are from its README in the llama.cpp repository, fetched 2026-09-29.
The two are not rivals in the usual sense. jev compiles in llama.cpp's tokenizer, and the jevos-v2 release also ships the model as GGUF files that llama-server can load, so the question is how much of the layer above the model you want to own.
This page is what llama-server provides, the gap to a decision endpoint, a sketch of filling it
yourself, the extras jev serve adds, and when llama-server is the better choice.
Its README calls it a "fast, lightweight, pure C/C++ HTTP server based on httplib, nlohmann::json
and llama.cpp." The feature list includes OpenAI-compatible chat completion, completion and
embedding routes, an Anthropic Messages compatible route, a reranking endpoint, parallel decoding
with multiple users through slots, grammar and JSON schema constraints, prompt caching, and token
probabilities. It listens on 127.0.0.1:8080 by default, takes --api-key for authentication, and
exposes /health and /tokenize.
For classification, three options on /completion matter most:
-
n_predict: the maximum number of tokens to generate. The README notes that at 0 "no tokens will be generated but the prompt is evaluated into the cache." -
n_probs: "If greater than 0, the response also contains the probabilities of top N tokens for each generated token." -
cache_prompt: re-use the KV cache from a previous request so that "the common prefix does not have to be re-processed."
There are also grammar and json_schema for constraining generation, and logit_bias for
nudging specific tokens.
To turn llama-server into a yes/no service you would write:
- A prompt that frames the text and the question for your chosen model, including its chat template, and makes a one-token answer the natural continuation.
-
A probability reader. With
n_predictat 1 andn_probsset, you get the top alternatives for the first generated token. You then find which of them mean yes and which mean no, across tokenizer variants such as "Yes", " yes" and "YES", and turn their probabilities into one P(yes). If neither appears in the top N, you need a policy for that. -
A request format for several named questions about one text, and code that sends them so
the shared text benefits from
cache_prompt. - Validation and errors: what a missing field, an over-long text or an unknown question type returns.
- Calibration checks on your own labelled cases, because a general chat model's token probabilities for "yes" are not necessarily calibrated for your questions.
None of this is hard in isolation. Together it is a small project, and each piece has edge cases.
An untested sketch, to show the shape, not a recipe. It asks llama-server for one token with its top alternatives:
curl http://127.0.0.1:8080/completion -H 'Content-Type: application/json' -d '{
"prompt": "<your template around the text and the question>",
"n_predict": 1,
"n_probs": 10,
"cache_prompt": true
}'The completion_probabilities field of the response holds, for the generated token, a
top_logprobs list of up to n_probs entries with the token text and its log probability. From
there, step 2 above is your code. The general reason this route still costs a decode step is on
why one forward pass beats generating an answer.
jev serve --threads 16 gives you, on 127.0.0.1:8017:
-
One endpoint, one answer shape.
state(text or any JSON) plus named questions in; each yes/no question back as{"type": "noul", "noul": ...}, eachchoicequestion as the most probable option with a probability per option, eachscorequestion as the expected level with a probability per level, withoutput_tokensalways 0. - Shared reading of the text. Questions in one request share the state, which is read once: three questions took about 66 ms against 49 ms for one on our reference laptop.
-
A wire format someone else defined. It is TypeSafe Jev's, so clients written for Jev's SDK
work unchanged for yes/no,
choiceandscorequestions, as described on an open-source alternative to Jev. -
Provenance.
GET /healthreports the served model; every successful response carries aServer-Timingheader. - A model built for yes/no questions, not a general chat model prompted into answering them, with calibration measured on held-out questions (0.009 on 6,397 natural yes/no questions).
-
Server-free batch use.
jev decideanswers a request file, the same body asPOST /v1/systemone, with no server, and prints the answers as indented JSON.
What it does not add: generation, chat, embeddings or other models. score questions are
answered, but early: the most probable level is right 54% of the time on 2,350 held-out
questions, and weak where the level is a sum of points.
- You want a different or larger model, or several.
- You need generation, embeddings or reranking from the same process.
- You need parallel slots for many concurrent users; jev reads small requests arriving together in one model call and reached 10.1 requests/s with 8 clients on our reference laptop (median 780 ms), a capacity figure, not a latency one, a distinction made on throughput vs latency for a decision server.
- You want to own every line of the prompt and the probability logic.
Can llama.cpp server do classification? Yes, with your own prompt and code that reads token
probabilities from n_probs. It gives the parts, not the classifier.
Does jev serve use llama-server? No. It is one native binary that runs jevos-v2 with 8-bit weights through OpenVINO, uses llama.cpp only as its tokenizer, and exposes its own endpoint.
How do I get a yes/no probability from llama-server? Generate one token with n_probs set,
then add up the probabilities of the tokens that mean yes and those that mean no, and normalise.
Which one is faster? We have not measured llama-server with a comparable setup. jevos took 26 and 112 ms on our short and long requests, read from scratch, on an Intel Core Ultra 7 255H.
Can I run both? Yes; they default to different ports, 8080 and 8017.
See also: llama-cpp-python vs ctypes, using llama.cpp prebuilt binaries and curl examples for a local LLM decision API.
- llama-server description, features, default host and port,
--api-key,/health,/tokenize,n_predict,n_probs,cache_prompt,grammar,json_schema,logit_biasandcompletion_probabilities: the llama.cpp server README, fetched 2026-09-29. -
jev serve,jev decide, the endpoints,Server-Timingand the 422: the jev README; latency figures are our own measurements. - Calibration error 0.009: our held-out split, 6,397 natural yes/no questions,
jevos-q8_0.
From the notes of jev, which borrows llama.cpp's tokenizer and says so.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building