-
Notifications
You must be signed in to change notification settings - Fork 132
securing a local llm server with an api key
Set the JEV_API_KEY environment variable when you start jev serve, and every call except
GET /health then needs Authorization: Bearer <key>; anything else gets a 401. Keep the
default bind address, 127.0.0.1, unless another machine has to reach the server, and if one
does, put a reverse proxy with TLS in front, because jev serve itself speaks plain HTTP and the
key would otherwise cross the network in clear text.
Without a key, the server trusts anyone who can reach its port. On 127.0.0.1 that means
programs on the same machine, which is often fine. The moment you pass --host to listen on a
network address, the key stops being optional.
This page is how to turn the key on and check it, what remains reachable without it, the bind address, TLS through a proxy, and what the key does not protect against. The commands are minimal sketches to adapt; the behaviour described is read from jev's source.
The variable is read once, when the server starts:
export JEV_API_KEY="$(openssl rand -hex 32)"
./jev serve$env:JEV_API_KEY = "<a long random string>"
.\jev.exe serveThe start-up line on stderr says which mode you are in: it ends with (Bearer auth) when a key
is set and (no auth) when it is not. Check that line after every deployment change; it is the
fastest way to catch a service manager that did not pass the variable through. An empty value
counts as no key.
Clients send the header on every request:
curl http://127.0.0.1:8017/v1/models -H "Authorization: Bearer $JEV_API_KEY"A missing or wrong key gets 401 with WWW-Authenticate: Bearer and the detail "Missing or
invalid API key". The server compares the whole header, Bearer plus the key. This is the same Bearer
scheme the hosted Jev API uses, which is why clients written for it work unchanged; how to
switch such a client is on an open-source alternative to Jev.
Every call except GET /health checks the key. With a key set:
| Path | Needs the key | What it exposes |
|---|---|---|
POST /v1/systemone |
yes | the decisions |
GET /v1/models |
yes | served model name and the jev-latest alias |
GET /health |
no, by design | readiness and the model name |
/health stays open so that load balancers and process supervisors can check readiness without
a secret. It tells a caller which model you run, not what you asked or were answered. If even
that is too much on your network, block the path at the proxy.
--host defaults to 127.0.0.1 and --port to 8017. On the loopback address the server is
unreachable from other machines whatever the key says, which is the safest setting and the right
one when the client, a script, n8n or Home Assistant, runs on the same machine. Clients in a
container are the usual reason to change it; the details for that case are on
local LLM yes/no decisions in n8n.
If you listen on a network address, do three things together: set the key, restrict the port to the machines that need it with your firewall, and add TLS.
jev serve has no certificate options, so it serves HTTP only. For anything that leaves
the machine, terminate TLS in a reverse proxy and keep jev on 127.0.0.1 behind it. With Caddy,
whose documentation says it "will serve your proxy over HTTPS automatically and by default if it
knows the hostname", a two-line Caddyfile is enough:
decisions.example.com
reverse_proxy 127.0.0.1:8017
Per Caddy's docs, a public domain needs DNS pointing at the machine and ports 80 and 443 open to
get a publicly trusted certificate; .localhost names get self-signed ones. Any other proxy you
already run works the same way. The proxy is also where to add what jev does not have: rate
limits, IP allow-lists and access logs. The operational side of running your own model server is
on self-hosted AI for decisions.
Be clear about the scope:
- One key, one role. There are no users, scopes or per-client keys. Anyone with the key can ask anything. Changing it means restarting the server with a new value.
- No rate limiting. The server runs one model on the CPU: small requests arriving together are read in one model call, but throughput stays around 10 requests per second on the reference laptop, so a single client sending a flood keeps everyone else waiting. A request is bounded (the state at 256 KB, at most 1,024 questions, the context at 8,192 tokens by default), but the number of requests is not.
- Not a data policy. The key controls who can call the server; it says nothing about who can read your logs or the texts your application stores. That part is on a private LLM for text classification.
- Never in a browser. A key in front-end code is public. Call the server from your backend; calling a local LLM decision server from JavaScript explains why the browser cannot call it cross-origin anyway.
How do I add an API key to jev serve? Set JEV_API_KEY in the environment before starting
it. Clients then send Authorization: Bearer <key>.
Why does /health work without the key? On purpose: readiness checks should not need a secret. It reports readiness and the model name, not decisions.
Does jev serve support HTTPS? Not by itself. Put a reverse proxy with TLS in front and keep
the server on 127.0.0.1.
Is it safe to expose on the internet? Not on its own. With a key, TLS, a firewall rule and rate limits in a proxy it is a normal internal service; without them it is not.
See also: curl examples for a local LLM decision API, on-premise LLM for business decisions and logging LLM decisions for audit.
-
JEV_API_KEY, the start-up line, the 401 response, the header comparison, which routes check the key, the--host/--portdefaults, plain HTTP and the request bounds: read from the source of jev. Throughput: our measurements on an Intel Core Ultra 7 255H laptop. - Caddy: reverse proxy quick-start, fetched 2026-09-29.
From the notes of jev, a yes/no decision model that runs on a laptop CPU. The start-up line prints "no auth" on purpose: it should be the first thing you see.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building