-
Notifications
You must be signed in to change notification settings - Fork 132
llm decisions vs keyword rules
Use regex and keyword rules for things with a fixed shape, such as an order number, an email address, a URL or an exact phrase, and a yes/no model for things defined by meaning, such as "is the customer asking for a refund?". A regex is exact, practically free and fully predictable, and it breaks the moment someone phrases the same idea differently or negates it. A model handles paraphrase and negation, costs tens to hundreds of milliseconds per text on a CPU, and returns a probability instead of a certainty. The best systems use both: regex to extract, the model to judge.
Conflict of interest, in one line: we build jevos, a yes/no model, and this page still tells you to keep your regexes for everything they do well.
The failure of keyword rules is not that they are wrong, it is that they are wrong silently. A list that catches "refund" and "money back" misses "I want what I paid returned to my card", and nothing tells you it missed. A model's failures show up as uncertain probabilities you can route to review.
This page is what regex does best, where keywords break, what a model adds, a hybrid design, the costs, and a note on regex that can hurt you.
Anything with a syntax. Order IDs, dates in a known format, currency amounts, email addresses, URLs, SKUs, error codes in logs. Here a pattern is the right tool on every axis: exact, auditable, testable with a few examples, and orders of magnitude cheaper than any model. Even OpenAI's latency guide has a section titled "Don't default to an LLM", pointing out that language models "are therefore sometimes used in cases where a faster classical method would be more appropriate."
If the question is "does this message contain an order number?", write the regex.
Keywords approximate meaning with strings, and meaning escapes in three common ways.
Paraphrase. "Cancel my subscription", "I don't want to be billed next month", "how do I stop this". One intent, no shared keyword. Each new phrasing means a new rule, and the list is never finished.
Negation. "I am not asking for a refund, just an explanation" contains the keyword and means the opposite. Keyword lists that try to handle "not" within a few words of the keyword become the most fragile part of the system.
Context. "This is sick" in a gaming chat and in a medical form. "Kill the process" in a server log and in a threat. A string match has no idea which one it is looking at.
A question states the meaning directly and lets the model find it in the text:
{
"model": "jev-latest",
"state": "I'm not asking for a refund, I just want to know why I was charged twice.",
"questions": {
"wants_refund": {"type": "noul", "instructions": "Is the customer asking for their money back?"},
"billing_issue": {"type": "noul", "instructions": "Is this a billing problem?"}
}
}Each question comes back as its own noul, the probability of yes. On our 999 questions written
after training, jevos was right 0.893 of the time on paraphrase questions, 0.858 on negation
questions and 0.954 on facts stated in the text, the three cases where keyword rules are weakest.
What it does not add is exactness. A probability of 0.8 is a strong yes, not a proof; the model reads English only; and on arithmetic, 0.584 in the same test, it is worse than a line of code. Calculations belong in code, as small LLMs and arithmetic in yes/no questions shows.
The two work best in sequence, each doing its own part:
- Regex extracts the structured parts. Order IDs, amounts, dates, links. These become fields.
- Code computes what depends on them. Days since the order, whether the amount is over a limit, whether a link's domain is on your list.
-
The model answers questions of meaning about the text, with the computed fields in
statewhen they help: "Is the customer asking for their money back?", "Is the tone threatening?". - Code combines the answers with thresholds and rules, and sends the uncertain middle to a person.
A cheap regex can also act as a pre-filter where a pattern is decisive, such as skipping the model for auto-generated notifications. And an old keyword list is valuable as a test set: every text it matched, reviewed by hand, is a labelled case for checking the model. The pattern for phishing, where header checks and link parsing are exact and "does this ask for credentials?" is a meaning question, is on phishing email screening with a local LLM.
| regex / keywords | yes/no model (jevos) | |
|---|---|---|
| Cost per text | negligible | 26 ms (about 30 tokens) to 112 ms (about 190) on our laptop CPU |
| Memory | negligible | about 1 GB with the model loaded |
| Handles paraphrase | only what you listed | yes, 0.893 in our test |
| Handles negation | badly, by special cases | 0.858 in our test |
| Exact on syntax | yes | no, and not needed |
| Output | match or no match | P(yes) per question |
| Changing a rule | edit a pattern | edit a question |
Several questions on the same text cost much less than separate calls, because the text is read once: three questions took about 66 ms against 49 ms for one. So the model is affordable for most message streams on a CPU, but not free, and nothing beats a regex on price.
Regex has its own failure mode that models do not: catastrophic backtracking. OWASP describes
ReDoS as an attack that exploits the fact that most regex implementations "may reach extreme
situations that cause them to work very slowly (exponentially related to input size)." Its example
is ^(a+)+$, where each extra "a" in a failing input doubles the number of paths the engine tries.
If your patterns run on user text, review them for nested quantifiers, or use an engine that does
not backtrack.
Is an LLM better than regex for text classification? For meaning, such as intent, tone and paraphrase, usually yes. For fixed formats like IDs and emails, regex is better and far cheaper.
Why do keyword filters miss things? People phrase the same idea many ways, negate keywords and use words in more than one sense. A list cannot enumerate meaning.
Can I combine regex and an LLM? Yes: regex to extract structured fields, code to compute, the model to answer meaning questions, code to combine.
How does a yes/no model handle negation? Better than keyword rules: 0.858 on negation questions in our test. Positive phrasing of the question still helps.
Is a model fast enough to replace keyword rules? For most message streams on a CPU, yes: 26 to 112 ms per request in our measurement. It will never be as cheap as a regex.
See also: negation in yes/no questions, checking text for personal data with yes/no questions and a yes/no LLM vs a business rules engine.
- Accuracy by kind (paraphrase 0.893, negation 0.858, fact 0.954, arithmetic 0.584): our 999-question
test set on
jevos-q4_k_m. - Latency, memory and the three-question timing: our own measurements, in the jev README.
- "Don't default to an LLM": OpenAI, Latency optimization, fetched 2026-09-29.
- ReDoS definition and the
^(a+)+$example: OWASP, Regular expression Denial of Service, fetched 2026-09-29.
From the notes of jev, a yes/no model for the questions a pattern cannot express, and no use at all for finding an order number.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building