-
Notifications
You must be signed in to change notification settings - Fork 131
jevos vs llama guard
For the standard safety categories, violent crime, self-harm, child exploitation, hate and the rest of a published hazard list, a dedicated safety model such as Meta's Llama Guard is the specialist, and it is the tool to use. jevos does something different: it answers yes/no questions you write, so it fits the rules that are yours, such as "is this post off-topic for a cooking forum?", "does this reply promise a refund?" or "does the message mention a competitor?". Llama Guard 3 covers eight languages; jevos reads English only. Many systems need both: a safety classifier for the hazards everyone shares, and questions for the policies only you have.
Conflict of interest, in one line: we build jevos; the Llama Guard facts come from Meta's model card on Hugging Face, fetched 2026-09-29.
The confusion comes from both being "a model that says whether content is acceptable". The question that separates them is who wrote the definition of acceptable: a published taxonomy, or you.
This page is what Llama Guard is built for, what jevos does instead, where the specialist wins, where your own questions fit, how to use both, and the limits they share.
Meta's card describes Llama Guard 3 as "a Llama-3.1-8B pretrained model, fine-tuned for content safety classification." It can classify both sides of a conversation: "LLM inputs (prompt classification)" and "LLM responses (response classification)." Its output is generated text, saying whether the content is safe or unsafe and, when unsafe, listing the violated categories.
The categories are fixed. The card lists 14, based on the MLCommons taxonomy of 13 hazards plus one for code interpreter abuse:
S1 violent crimes, S2 non-violent crimes, S3 sex-related crimes, S4 child sexual exploitation, S5 defamation, S6 specialized advice, S7 privacy, S8 intellectual property, S9 indiscriminate weapons, S10 hate, S11 suicide and self-harm, S12 sexual content, S13 elections, S14 code interpreter abuse.
For a score instead of a label, the card says: "We look at the probability for the first token, and use that as the 'unsafe' class probability. We can then apply score thresholding to make binary decisions." It is 8B parameters in bfloat16, with an INT8 version listed, under the Llama 3.1 Community License.
jevos has no taxonomy. You send a text as state and any number of named yes/no questions, and
each comes back as its own noul, P(yes):
{
"model": "jev-latest",
"state": "Check out my channel for better recipes than this one, link in bio",
"questions": {
"off_topic": {"type": "noul", "instructions": "Is this comment unrelated to the recipe it was posted under?"},
"self_promo": {"type": "noul", "instructions": "Does the comment promote the writer's own channel or product?"},
"insulting": {"type": "noul", "instructions": "Does the comment insult another person?"}
}
}None of those three is a safety hazard in a published taxonomy, and all three are the kind of rule a community actually enforces. That is the space jevos is for. It is a 1B-class model on a CPU: 26 to 112 ms per request on our reference laptop, with the text read once for every question in the request.
Say it plainly: for the categories in its taxonomy, use a safety model trained for them.
- Coverage of hard categories. Child exploitation, weapons, self-harm: these need a model built and evaluated for them, not a general yes/no reader asked a question on the fly.
- Languages. The card lists English, French, German, Hindi, Italian, Portuguese, Spanish and Thai. jevos is English only.
- Conversation roles. Llama Guard is built to classify prompts and responses in a chat.
- A shared vocabulary. Categories with codes make reporting and audits comparable across systems.
Asking jevos "is this content sexual exploitation of a minor?" is the wrong design. The right one is a specialist model plus the legal and reporting process that category requires.
Most moderation work is not in the hazard list. It is forum rules, brand rules, product rules: off-topic posts, self-promotion, personal data in public reviews, a reply that promises what support cannot give, a message asking to cancel. These change often, differ per community, and are easy to state as a question. jevos' measured strengths match them: 0.938 on tone, 0.859 on intent and 0.954 on facts stated in the text, in our 999-question test.
The operating pattern is on content moderation with a local LLM: act on the confident ends, and send the middle band to a moderator.
A layered setup is straightforward:
- Run the safety classifier on every item. Anything it flags follows your safety process.
- Ask your own yes/no questions on the rest: the community and product rules.
- Route uncertain answers from either to a person, and log both the scores and the decision.
Keep the two apart in logs and in code. A "safe" verdict from a safety model does not mean the post follows your rules, and a low P(yes) on your questions does not mean the post is safe.
Meta's card is candid about Llama Guard's limits: some categories "may require factual, up-to-date knowledge to be evaluated", naming defamation, intellectual property and elections, and the model is susceptible to adversarial and prompt injection attacks. jevos has the same exposure to adversarial text, and a smaller model has less knowledge to draw on. Neither is a security boundary on its own; see prompt injection screening with a small model for how far a screen goes.
What is Llama Guard? Meta's content safety classifier: Llama Guard 3 is an 8B model that labels prompts and responses safe or unsafe against 14 hazard categories.
Can jevos replace Llama Guard? No, not for the hazards in its taxonomy. Use a safety model there, and jevos for your own policy questions.
Can Llama Guard give a probability? The card describes using the first token's probability as the unsafe score and thresholding it.
Which languages? Llama Guard 3 lists eight; jevos reads English only.
Is either a security control? No. Both can be fooled by adversarial input and belong in a layered design.
See also: AI agent guardrails with yes/no questions, review moderation with a local LLM and jevos vs Jev vs Laya.
- Everything about Llama Guard 3 (base model, size, categories, languages, first-token probability, limits, licence, INT8 version): Meta's model card on Hugging Face, fetched 2026-09-29.
- jevos latency and accuracy by kind: our own measurements, see the
jev README and our 999-question test set on
jevos-q4_k_m.
From the notes of jev, which leaves the hazard list to the models built for it.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building