-
Notifications
You must be signed in to change notification settings - Fork 135
what is gguf
GGUF is a single binary file that holds a model's weights together with everything needed to load and run it: the architecture name, the tokenizer, the prompt template and the quantization type. It is the native format of llama.cpp and of the engines built on ggml. For deployment that means the model is one artifact: you copy one file, check one hash, and point a runtime at it. No folder of configs and tokenizer files has to travel with it.
The non-obvious point for a classifier is that the file is also the model's identity. If the hash of the file is the same, the weights, the tokenizer and the template are the same, so the same request gives the same answer on the same runtime. That makes the file hash the thing to log, to pin in a deployment and to check after a download.
This page is what is inside a GGUF file, why the single-file design matters when you ship a decision model, which metadata is worth reading, where GGUF files come from, and what the format does not tell you.
The format is specified in the ggml repository. A file is laid out in order:
- a magic number (the bytes "GGUF") and a format version, currently 3;
- the number of tensors and a list of metadata key-value pairs;
- one info record per tensor: name, dimensions, type and offset;
- padding to an alignment boundary (32 bytes unless the file says otherwise);
- the tensor data itself.
The spec lists its own goals plainly: single-file deployment ("do not require any external files"), extensibility without breaking old readers, mmap compatibility so models load fast, ease of use without external libraries, and "all information needed to load a model is contained in the model file." GGUF replaced three earlier ggml formats (GGML, GGMF and GGJT), which lacked versioning, a way to name the architecture, or both.
A classifier in production is usually a small service that must be reproducible and boring. Three properties of GGUF help directly.
One artifact to verify. The jevos-v4 release includes the model as two GGUF files, for
llama.cpp and tools built on it: jevos-v4-q4_k_m.gguf and jevos-v4-q8_0.gguf. Computing the
sha256 of one file and comparing it with a copy you trust tells you the whole model arrived intact, which is what makes
offline and air-gapped use practical.
One artifact to identify. One sha256 covers the weights, the tokenizer and the template together. If you run a GGUF and log decisions for audit, that hash is the field that says which model produced a probability.
Fast, cheap loading. Because tensors sit at aligned offsets, a runtime can memory-map the file instead of parsing and copying it. llama.cpp's default load mode maps the model unless the device does not support it.
Most of what a deployer wants to know is in the key-value section, before any weights:
| Key | What it tells you |
|---|---|
general.architecture |
which model family the runtime must implement |
general.file_type |
the overall quantization of the file |
general.quantization_version |
required when tensors are quantized |
general.name, general.version
|
human labels set by whoever converted it |
tokenizer.ggml.model |
the tokenizer type |
tokenizer.chat_template |
the Jinja template that turns messages into a prompt |
You do not need to write a parser to see them. The Hugging Face Hub has a viewer that shows a
GGUF file's metadata and tensor list on the model page, and the @huggingface/gguf JavaScript
package reads the header of a remotely hosted file without downloading the weights.
jev itself does not read GGUF files: it runs the same model with 8-bit (INT8) weights through
OpenVINO, from the model folder beside the binary. Its tokenizer is llama.cpp's, compiled in,
so its token ids match the GGUF files.
Models are usually developed in PyTorch and stored as safetensors. The llama.cpp repository
ships convert_hf_to_gguf.py to turn such a checkpoint into GGUF, and the llama-quantize tool
to turn a high-precision GGUF (F32 or BF16) into a smaller quantized one. The quantization names
you see on file names, such as Q4_K_M, are explained on
GGUF quantization types explained.
On the reading side, the Hugging Face docs list llama.cpp, LM Studio, GPT4All and Ollama as tools that use GGUF. Ollama's import guide takes a GGUF path in a Modelfile and says plainly that Ollama does not quantize GGUF files during import. What Ollama adds on top of llama.cpp, and whether a classification service needs it, is on llama.cpp vs Ollama for a classification service. Transformers can also load a GGUF, but on most devices it dequantizes the weights at load time into an ordinary dense model, which gives up the size advantage. The comparison with the other common format is on GGUF vs safetensors.
Being straight about the limit: GGUF is a container, not a certificate.
- It does not say how well the model does your task. Two files with the same metadata can answer your questions very differently. Only a test set built from your own cases tells you that.
-
The file name is a convention. The spec recommends a naming pattern (base name, size
label, version, encoding) but nothing enforces it. Read
general.file_typeand the hash, not the name. - It does not pin the runtime. A new llama.cpp release can read the same file with different kernels. For reproducible answers, pin the runtime release as well as the file.
Is GGUF only for llama.cpp? It is llama.cpp's native format, and the Hugging Face docs list other local tools that use it, including Ollama and LM Studio. Transformers can load it too, usually by dequantizing it.
Can I use my own GGUF with jev? No. jev does not read GGUF files. To run a GGUF, your own or the two jevos-v4 release files, use llama.cpp or a tool built on it directly.
How do I check a downloaded GGUF is intact? Compute its sha256 and compare it with a copy you trust, for instance the hash you recorded when you first checked the download.
Does a GGUF file contain the tokenizer? Yes. The tokenizer and, usually, the chat template are metadata keys inside the file.
What does the version number in the header mean? It is the format version, currently 3. It
says nothing about the model's own version, which lives in general.version if set.
See also: GGUF quantization types explained, using llama.cpp prebuilt binaries and self-hosted AI for decisions.
- Our own facts: the release file names, from the jevos release; the INT8 model and the llama.cpp tokenizer, from the jev source code.
- GGUF specification, ggml repository, fetched 2026-09-29: layout, goals, metadata keys, naming convention, predecessor formats.
-
Hugging Face Hub docs: GGUF, fetched 2026-09-29:
metadata viewer,
@huggingface/gguf, tools that use GGUF. - llama-quantize README and llama.cpp completion README (load modes), fetched 2026-09-29.
- Ollama import guide and Transformers GGUF docs, fetched 2026-09-29.
From the notes of jev, which ships its model as an OpenVINO build for its own binary and, for other tools, as two GGUF files.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building