-
Notifications
You must be signed in to change notification settings - Fork 137
throughput vs latency for a decision server
Latency is how long one decision takes; throughput is how many decisions a server completes per second, and one does not tell you the other. Our published jevos figures, 28 ms for a short request and 130 ms for a long one on a laptop CPU, are latency: one request at a time, on an idle machine. They say nothing about what happens when a hundred requests arrive together, which depends on how the server schedules work, how long the queue gets, and how many copies of it you run.
The link between the two is waiting. When requests arrive faster than they are served, latency stops being the model's time and becomes the model's time plus the queue. A server can have excellent latency at low load and poor latency at peak without anything about the model changing.
This page is the two definitions, the one law that connects them, how jev serve handles concurrent requests, the kinds of batching, and how to measure capacity for your own traffic. The one concurrency measurement here is from one laptop and from an earlier jevos version (it was not repeated on jevos-v4): a starting point for yours, not a capacity figure.
| Latency | Throughput | |
|---|---|---|
| Question it answers | How long does my request wait? | How much traffic can this handle? |
| Unit | milliseconds per request | requests (or decisions) per second |
| Measured with | one request at a time, repeated | many concurrent requests, sustained |
| Who needs it | the user or caller waiting | whoever sizes the servers |
NVIDIA's benchmarking documentation defines requests per second as completed requests divided by the duration of the test, measured across all simultaneous requests. That is a property of the whole system under load, not of a single request.
Little's law says that in a stable system the average number of items inside equals the arrival rate times the average time each spends inside: L = lambda x W. It holds whatever the arrival pattern or service order.
An illustrative example, with made-up traffic, to show how it works: if decisions arrive at 20 per second and each spends 0.25 s in the system, then on average 5 are in the system at any moment. If the server can only work on one at a time and each takes 0.2 s, it can finish at most 5 per second, and 20 arriving per second means the queue grows without limit. Latency in that situation is not 0.2 s; it is however long the queue has become.
The practical reading: a latency figure is only meaningful up to the load where the queue stays short. Past that, more traffic means more waiting, not more throughput.
jev serve reads small requests that arrive at the same time together, in one model call, up to
--batch-tokens (384 tokens by default). Past that, requests wait for the model. The time a
request waits is inside the total duration of the Server-Timing header, while inference
is only the model's work. So under load, a growing gap between total and inference is the
queue, visible per response.
We measured it with one-question requests from our 999-question set, on the reference laptop (Intel Core Ultra 7 255H, 16 threads), with an earlier jevos version; the numbers for jevos-v4 are not published, and its single-request latency is the 28 ms and 130 ms above:
| Concurrent clients | Requests per second | Median latency |
|---|---|---|
| 1 | 8.7 | 110 ms |
| 4 | 9.8 | 390 ms |
| 8 | 10.1 | 780 ms |
Throughput barely moves past one client while latency grows with the queue, which is Little's law on a server that is already busy. That is one laptop and one kind of request: real traffic mixes sizes, and your CPU is not ours.
To go beyond one process's ceiling, the options are the usual ones: more processes on a machine with cores to spare, more machines behind a load balancer, or a different serving stack. We measured about 1 GB of extra memory with the model loaded in one process; until you have checked how several processes share memory on your system, plan that much for each.
Questions in one request. Several questions about the same text share its reading. Three questions took about 66 ms together against 49 ms for one alone. This is the batching you control from the client, and it raises decisions per second without any server change; see many questions about one text.
Requests batched on the server. General serving systems process many users' requests together. The vLLM paper opens with "high throughput serving of large language models (LLMs) requires batching sufficiently many requests at a time", and the llama.cpp server lists "continuous batching" and parallel slots among its features. This raises throughput, usually at some cost to each request's latency, and it is where GPUs shine; see CPU or GPU for a small LLM. jev serve does a small form of this: small requests that arrive at the same time are read together in one model call.
Batching over files. For offline work, the unit is a file of requests and the measure is how
long the job takes. jev decide answers request files without a server; the workflow is on
batch decisions from files with jev decide.
- Get the single-request baseline first: median and p90 of your real requests, one at a time, as on measuring LLM latency.
- Add load in steps. Send requests from several concurrent clients, starting at a rate well below the ceiling and raising it.
-
Watch p90 and the queue, not just the average. Record the
totalminusinferencegap fromServer-Timingon every successful response. - Stop at the knee. Capacity is the highest rate at which the p90 still fits your budget, not the rate at which the server stops failing.
- Leave headroom. Traffic has peaks; size for them, not for the daily average.
Run the load generator on a different machine from the server if you can. A load generator on the same CPU competes with the model and lowers both numbers.
For a user waiting on a form, or an agent waiting before its next step, latency. For a nightly job over a million records, throughput. For a webhook that fans out to many decisions, both: the p90 has to fit the sender's deadline at the peak rate. The budgets for each case are on latency budgets: where a 200 ms model fits.
What is the difference between throughput and latency? Latency is the time for one request; throughput is how many requests a system completes per second under load.
Does low latency mean high throughput? Not by itself. It sets a ceiling when requests are served one at a time; batching and more copies raise throughput beyond it.
How many requests per second can jevos handle? With an earlier jevos version on our laptop, one process answered about 9 to 10 one-question requests per second, with the median latency growing from 110 ms at one client to 780 ms at eight; not measured on jevos-v4. Measure with your traffic.
Does jev serve process requests in parallel? Partly. Small requests that arrive together are
read in one model call; the others wait, and the wait is included in the total duration of
Server-Timing.
How do I increase throughput? Group questions per text, trim inputs, and run more processes or machines once one is saturated.
See also: the fastest AI model for yes/no decisions, prefill vs decode: where LLM latency comes from and self-hosted AI for decisions.
- Single-request latencies, the three-question timing, the concurrent-client measurement
(earlier jevos version), memory, and the
Server-Timingheader: our measurements and the jev README. Requests read together and queue time counted intotal: the jev source. - Little's law: Little's law, Wikipedia, fetched 2026-09-29.
- Requests per second definition: NVIDIA NIM benchmarking metrics, fetched 2026-09-29.
- Batching for throughput: Kwon et al., PagedAttention, and the llama.cpp server README, both fetched 2026-09-29.
- The traffic in the Little's law example is illustrative arithmetic, not a measurement.
From the notes of jev, whose server says in its timing header how long each request waited, which is where capacity planning should start.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building