-
Notifications
You must be signed in to change notification settings - Fork 132
intent detection with a local llm
For intent detection with a yes/no model, each intent in your catalogue is a question ("Does the user want to change their delivery address?"), all of them go in one request, and code picks the intents whose P(yes) clears a threshold, or hands off when none does. Adding an intent means adding a question, with no retraining and no example utterances to collect. Intent questions scored 0.859 on our test set of 999 new yes/no questions, which is good enough for routing and suggestions, and not good enough to trigger irreversible actions without a check.
The shift from a trained intent classifier is that intents stop competing. A classifier trained on ten intents must put every message in one of them; independent questions can say that a message has two intents, or none, and both answers are useful in a chat assistant.
This page is the catalogue as a request, adding intents, overlapping intents, the no-match fallback, what the measured number covers, how far a catalogue scales, and where running locally matters.
{
"model": "jev-latest",
"state": {
"channel": "in-app chat",
"previous_bot_message": "Your order 7731 is out for delivery today.",
"user_message": "ah no, I won't be home, can you leave it with a neighbour or send it to my office instead?"
},
"questions": {
"change_address": {"type": "noul", "instructions": "Does the user ask to deliver the order to a different address?"},
"delivery_instructions": {"type": "noul", "instructions": "Does the user ask to leave the parcel with someone or in a specific place?"},
"reschedule": {"type": "noul", "instructions": "Does the user ask to deliver the order on a different day?"},
"cancel_order": {"type": "noul", "instructions": "Does the user ask to cancel the order?"},
"track": {"type": "noul", "instructions": "Does the user ask where the order is or when it will arrive?"}
}
}Two habits carry most of the accuracy. Every question starts the same way ("Does the user ask
to ..."), which keeps the probabilities comparable with each other. And the previous bot message
is in state, because "send it to my office instead" only has an intent in the context of what
came before. How to phrase intent questions in general, including requests versus complaints,
is on asking about intent: what does the writer want?.
Adding "return an item" is one line in the questions. What it does not remove is the need to check that the new question does not take messages from the old ones. A cheap routine:
- Keep a file of a few hundred real messages with the intents a person assigned.
- Before shipping a new or reworded question, run the file through
jev decideor the server. - Compare, per intent, how many messages changed their top intent.
If "return an item" starts winning messages that were "cancel order", the two questions overlap, and the fix is in their wording ("return an item you already received" versus "cancel an order that has not shipped"). The same routine catches a question that was reworded for one case and broke three others; why wording changes an LLM's answer covers how to test paraphrases.
Sales teams apply the same catalogue idea to buying signals, such as budget, timeline and who decides, on lead qualification with yes/no questions.
The example message above has two plausible intents: change the address, or leave the parcel with a neighbour. A single-label classifier would pick one. With independent questions, code can do the right thing for a chat assistant:
def intents(p, act=0.7, ask=0.4):
sure = [k for k, v in p.items() if v >= act]
maybe = [k for k, v in p.items() if ask <= v < act]
if len(sure) == 1:
return {"do": sure[0]}
if len(sure) > 1:
return {"ask_user_to_choose": sorted(sure, key=p.get, reverse=True)}
if maybe:
return {"confirm": max(maybe, key=p.get)}
return {"handoff": True}Two confident intents become a short clarifying question to the user ("Change the address, or leave it with a neighbour?"). One uncertain intent becomes a confirmation ("Do you want to reschedule the delivery?"). The thresholds are placeholders to be set from your labelled messages.
The last line of that function is the one that matters most. Below the lower threshold on every intent, the message is outside your catalogue: small talk, a new kind of request, or a complaint with no ask. Hand it to a person or to a general reply, and log it. Those logs are your backlog of missing intents. A separate "other" question phrased as a negative ("Is this about none of the above?") is the weaker design, because it gives the model nothing in the text to find.
Intent scored 0.859 on 71 questions in our set of 999 written after training, and negation, which appears in intents such as "does not want a refund", 0.858. Those are measurements on our texts (emails, tickets, logs, reviews, forms), not on chat messages, and not on your catalogue. They are the reason to use thresholds with a confirm step rather than acting on the top intent alone.
The model reads stated intent better than implied intent. "Cancel my subscription" is stated. "I guess I won't be needing this anymore" implies it. If implied intents matter to you, ask for them explicitly ("Does the user suggest they may stop using the service?") as their own question, and route them to a person; the retention case is worked through on detecting cancellation intent in customer messages.
The 8,192-token context applies to the text plus each question, not to all the questions together, so a few dozen intents with a chat message fit easily. Cost grows with the total tokens read: on our reference laptop a 30-token request took 26 ms and a 191-token one 112 ms, and questions are tokens too. For a catalogue of hundreds of intents, ask in two stages: a handful of coarse questions first (orders, account, billing), then only the detailed intents under the winning area. The same idea, used for choosing between models, is on an LLM router with yes/no questions.
Intent detection sits inside a conversation turn, so its latency is added to every reply. On our reference laptop a short request takes about 26 ms, well inside a chat turn, while the hosted Jev API measured from Europe took about 344 ms on the same short request, mostly network. The user's messages, which often contain order numbers and addresses, also stay on your server.
Where local does not help: languages other than English, and assistants that need the model to also write the reply. jevos only decides; the reply comes from your templates or another model.
What is intent detection? Working out what a user wants from their message, such as cancelling an order or changing an address, so the system can act or route it.
Can I add intents without retraining? Yes. Each intent is a question; add one, then replay a labelled set of messages to check it does not overlap the others.
How do I handle messages with several intents? Keep every intent above the threshold, and ask the user to choose when more than one is confident.
How accurate is it? Intent questions scored 0.859 on our own test set. Measure on your own messages and catalogue.
See also: semantic routing vs yes/no questions, support ticket routing with yes/no questions and mainly about: questions for messages with several topics.
- Intent 0.859 (71 questions), negation 0.858: our 999-question test set,
jevos-q4_k_m. - Latency 26 ms short (30 tokens) and 112 ms long (191 tokens) read from scratch, hosted Jev 344 ms from Europe on the short request, 8,192-token context: the jev README and our reference-laptop measurements.
- Thresholds in the code are placeholders.
From the notes of jev. The most useful answer an intent detector gives is "none of these", because it is the list of intents you have not written yet.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building