-
Notifications
You must be signed in to change notification settings - Fork 132
why a hosted llm api cannot answer in 50 ms
A hosted LLM API usually cannot answer in 50 ms because the request has to cross the network and back before the answer exists, and on a fresh connection it crosses more than once. Light in fibre covers about 200 km per millisecond, a new HTTPS connection needs round trips for TCP and TLS before the request is even sent, and the provider may queue the request behind others. From our laptop in Europe, the hosted Jev API took 344 ms on a short request and 345 ms on a long one, while the same requests ran locally in 26 and 112 ms.
The flatness is the tell. When a request with six times more text takes the same time, the model is not what you are waiting for. That also means a hosted API can be the right choice for long texts or hard questions, where its fixed cost is spread over more work.
This page is the ledger of where the time goes, the physics nobody can optimise away, how to read our two numbers, what you can shave off, and when hosted is still the better option.
A call to a hosted API over HTTPS, from a client that has not talked to the server recently:
| Step | Cost | Can it be avoided? |
|---|---|---|
| DNS lookup | a round trip to a resolver, often cached | usually cached after the first call |
| TCP connection | one round trip (the three-way handshake) | reuse the connection |
| TLS 1.3 handshake | one more round trip for a full handshake | reuse the connection; 0-RTT resumption with caveats |
| Request and response | one round trip plus transfer time | no |
| Queueing at the provider | anything from nothing to seconds | not from your side |
| The model itself | prefill, plus decode if it writes | smaller model, shorter input |
RFC 9293 calls TCP's setup the "three-way (or three message) handshake". RFC 8446, which defines TLS 1.3, describes a full handshake completed in one round trip and adds a zero round-trip mode for resumed sessions, noting that its "security properties ... are weaker". So a cold HTTPS request costs about three round trips before the first byte of the answer, and a warm one on a reused connection costs one.
Signals in optical fibre travel at around 200,000 km per second, as the Wikipedia article on optical fibre puts it, about two thirds of the speed of light in a vacuum. That is 200 km per millisecond, one way. Its own example: a 16,000 km fibre path between Sydney and New York means a minimum delay of 80 ms.
For a 50 ms budget, that gives a hard ceiling on distance. A single round trip to a data centre 5,000 km away uses 50 ms in fibre alone, before routers, before TLS, before the model; real cable routes are not straight lines, so the practical distance is shorter. Nearby regions fit; an ocean away does not.
| Same two requests | short (about 30 tokens) | long (about 190 tokens) |
|---|---|---|
| jevos on the laptop, text read from scratch | 26 ms | 112 ms |
| Jev, hosted, from Europe, network included | 344 ms | 345 ms |
Locally, going from 30 to 190 tokens quadrupled the time, because the model reads every token. Hosted, it added one millisecond. The length-dependent part of the hosted call was lost in the fixed part: connection, round trips, and whatever happens before the provider's model starts.
What the numbers do not tell you is how that fixed part splits. We do not know where TypeSafe's servers are, and we measured from one location; from a client closer to them the floor would likely be lower. This is what one application in Europe sees, not a property of the service.
- Reuse connections. HTTP/1.1 "defaults to the use of persistent connections", per RFC 9112. A client that opens a new connection per call pays the TCP and TLS round trips every time; a pooled session pays them once. Check that your HTTP client keeps connections open between calls.
- Call from the right region. Put the code that calls the API close to the API. The cheapest round trip is a short one.
-
Batch questions. Several questions about one text in one request pay the network once.
With a Jev-compatible API, that is the same
questionsobject with more entries. - Move the decisions that need speed off the network. A local model has no round trip at all. The general trade-off is on local vs hosted LLM decisions: latency, cost, privacy.
What you cannot shave is the distance and the provider's queue. If your budget is 50 ms and the nearest region is far, no client-side change fits it.
- Accuracy matters more than milliseconds. On 2,000 policy questions none of them was tuned on, Jev was right 0.927 of the time against 0.810 for jevos. The comparison is on jevos vs Jev vs Laya.
- The texts are long. Local latency grows with every token read on our laptop; a hosted model's fixed cost matters less the more work each request carries.
- You need score answers past jevos's early ones, or languages other than English.
- The latency budget is seconds, not milliseconds. A nightly job or a webhook with a three-second timeout has room for a round trip; see latency budgets: where a 200 ms model fits.
Because jevos speaks Jev's wire format, the two can split the work: answer locally when the probability is confident, and send the uncertain middle to the hosted model with a change of URL. The pattern is on a model cascade: small model first, large model on doubt.
Why is my LLM API call slow even for a one-word answer? Much of the time is network and setup, not the model. On our measurement from Europe, a 30-token and a 190-token request took the same 344 to 345 ms.
Can any hosted API answer in 50 ms? Only from close by, on a reused connection, with little queueing and a fast model. The distance alone rules it out from far away.
Does streaming fix it? Streaming shows the first token sooner, but the first token still waits for the round trip and the prompt. For a yes/no decision you need the whole answer anyway.
How do I measure the network part? Compare wall-clock time on your side with any timing the
server reports; for a local jev server, the Server-Timing header gives the server's own
durations. See measuring LLM latency.
Is a local model always faster? No. It was on our two requests. On long documents, from a server near the provider, measure both.
See also: why LLM latency grows with the length of the text, the fastest AI model for yes/no decisions and jevos vs the OpenAI API for yes/no classification.
- Local and hosted latencies and the 2,000-question accuracy comparison: our own measurements, published in the jev README.
- TCP handshake: RFC 9293, fetched 2026-09-29.
- TLS 1.3 handshake and 0-RTT: RFC 8446, fetched 2026-09-29.
- Persistent connections: RFC 9112, section 9.3, fetched 2026-09-29.
- Signal speed in fibre and the Sydney to New York example: Optical fiber, Wikipedia, fetched 2026-09-29.
- Slack's three-second response requirement: Slack Events API, fetched 2026-09-29.
From the notes of jev, written on a laptop in Europe, which is where the 344 ms came from and the only place we can vouch for it.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building