-
Notifications
You must be signed in to change notification settings - Fork 129
gguf quantization types explained
A GGUF quantization name says how many bits each weight is stored in and which scheme packs
them: Q8_0 is 8-bit weights in blocks of 32 with one scale per block, Q4_K_M is a 4-bit
"K-quant" mix that keeps some tensors at 6 bits. The number is the nominal bit width, the
suffix after the underscore is the scheme (_0, _1, _K), and a trailing _S, _M or _L
names a mix of types across tensors. Names starting with IQ are newer types built around an
importance matrix.
The practical reading is that fewer bits mean a smaller file and less memory traffic per token, at some cost in accuracy that depends on the model and the task. The cost is not a property of the name: it has to be measured on what you actually ask the model.
This page is a decoder for the names, a table of real bits per weight, what a mix is, what an importance matrix does, and how the two jevos builds fit in.
Take the name apart from left to right.
-
Qand a number: quantized, with that many bits per weight before overhead.F16,BF16andF32are plain floating point, not quantized. -
_0and_1: the original "legacy" schemes. Weights are cut into blocks of 32;_0stores one scale per block (w = q * block_scale),_1stores a scale and a minimum (w = q * block_scale + block_minimum). Q4_0, Q4_1, Q5_0, Q5_1 and Q8_0 are all legacy. -
_K: the K-quants, added to llama.cpp in June 2023. Blocks are grouped into super-blocks and the per-block scales are themselves quantized, which spends fewer bits on overhead. Q4_K, for example, uses super-blocks of 8 blocks of 32 weights with 6-bit scales and minimums, for 4.5 bits per weight. -
_S,_M,_L: small, medium, large. These are not new storage types but recipes that decide which tensors get a higher-precision type. -
IQ: types such as IQ4_NL, IQ4_XS, IQ3_S and IQ2_XXS use super-blocks of 256 weights and are designed to be used with an importance matrix. They extend to much lower bit rates: IQ1_S is listed at 1.56 bits per weight, against 2.625 for Q2_K, the smallest K-quant.
There are also ternary types (TQ1_0, TQ2_0) and MXFP4, a 4-bit block floating point type, which you will meet less often.
The nominal number understates the file size, because scales and mixes add bits. The llama-quantize README publishes the effective figures for its 8B reference model:
| Type | Bits per weight | Size of the 8B reference |
|---|---|---|
| IQ4_XS | 4.46 | 4.17 GiB |
| Q4_K_S | 4.67 | 4.36 GiB |
| IQ4_NL | 4.68 | 4.38 GiB |
| Q4_K_M | 4.89 | 4.58 GiB |
| Q5_K_M | 5.70 | 5.33 GiB |
| Q6_K | 6.56 | 6.14 GiB |
| Q8_0 | 8.50 | 7.95 GiB |
| F16 | 16.00 | 14.96 GiB |
Q8_0 costs 8.5 bits, not 8, because each block of 32 weights carries its scale. Q4_K_M costs almost 4.9, not 4.5, because of the tensors the mix keeps at higher precision.
When the K-quants were introduced, the pull request described the two 4-bit recipes this way: Q4_K_S "uses GGML_TYPE_Q4_K for all tensors", while Q4_K_M "uses GGML_TYPE_Q6_K for half of the attention.wv and feed_forward.w2 tensors, else GGML_TYPE_Q4_K". The idea is that some weight matrices are more sensitive to rounding than others, so spending extra bits on them buys back more quality than spreading the same bits evenly. The recipe lives in llama.cpp's code and can change between releases, so read that 2023 description as the idea, not as a specification of every Q4_K_M file you download.
llama-quantize --imatrix file takes an importance matrix, data that tells the quantizer which
weights matter most, and uses it, in the README's words, "for quant optimizations". The README
puts the
general trade plainly: quantization "may introduce some accuracy loss which is usually measured
in Perplexity (ppl) and/or Kullback-Leibler Divergence (kld). This can be minimized by using a
suitable imatrix file." The IQ types rely on it; K-quants can use it too.
For a classifier, note what those two metrics measure: how close the quantized model's next-token predictions stay to the original on general text. They are a good sign, not a measurement of your yes/no accuracy.
jevos ships two GGUF files: jevos-v2-q4_k_m.gguf at 619 MB and jevos-v2-q8_0.gguf at 943 MB,
for llama.cpp, Ollama, LM Studio and other tools. jev itself reads neither: it runs the same model
with 8-bit (INT8) weights through OpenVINO, and on 215 parity cases its answers are within 0.056
of the q8_0 file's. The speed and size side of choosing between the two GGUF files has its own
page, Q4_K_M vs Q8_0 for a small model.
Being straight about the limit: the accuracy of q4_k_m and q8_0 has not been measured on the same question set. Our published numbers come from different sets on different builds, so they cannot be subtracted to give a quantization cost. If the difference matters to you, run both files on your own labelled cases.
- Start with Q4_K_M. It is the type in the llama-quantize README's basic example. On a CPU the smaller file is usually also the faster one, for reasons explained on what makes a local LLM fast on a CPU.
- Use Q8_0 when you want the reference. It is closest to the original weights, costs more memory and time, and is the natural baseline to test a smaller build against.
- Go below 4 bits only with a test set. IQ3 and IQ2 files are much smaller, and small models have the least spare capacity to lose. Measure on a yes/no test set of your own data before shipping one.
What does K mean in Q4_K_M? It marks the K-quants, a scheme with super-blocks and quantized block scales. The M means the medium mix, which keeps some tensors at a higher-precision type.
Is Q8_0 lossless? No. It is 8-bit round-to-nearest with one scale per 32 weights. It is close to the original, but it is still a quantization.
What is the difference between Q4_0 and Q4_K_M? Q4_0 is a legacy type with one scale per block of 32; Q4_K_M uses super-blocks, quantized block scales and a mix of types across tensors.
Do I need an imatrix? Not for Q8_0 or the common K-quants, though it can help. The IQ types are designed around one.
Which quantization is most accurate for classification? Nobody can say from the name. The higher-bit types are usually closer to the original model; measure the ones you consider.
See also: what is GGUF, GGUF vs safetensors and run an LLM locally without a GPU.
- Our own measurements: file sizes from the jevos release; the 215 parity cases comparing jev's 8-bit answers with the q8_0 file's; the statement that accuracy of the two GGUF builds was not compared on one set.
- llama-quantize README, fetched 2026-09-29: bits per weight and sizes, imatrix option, the accuracy-loss sentence.
- Hugging Face Hub docs: GGUF quantization types, fetched 2026-09-29: block and super-block structure and weight formulas.
- llama.cpp pull request 1684, k-quants, fetched 2026-09-29: Q4_K_S and Q4_K_M recipes as first described, June 2023.
From the notes of jev, which ships a 4-bit and an 8-bit build so that anyone can test the smaller one against the larger on their own questions.
- Ask a local LLM a yes/no question and get P(yes)
- Zero-shot text classification with yes/no questions
- LLM policy decisions: put the rule in the question
- LLM as a judge on a CPU
- Why a small LLM says yes when the answer is no
- Small LLMs and arithmetic in yes/no questions
- Our held-out benchmark said 0.855, new questions said 0.757
- jevos vs Jev vs Laya for yes/no decisions
- An open-source alternative to Jev for yes/no decisions
- jevos vs the OpenAI API for yes/no classification
- jevos vs Ollama for yes/no decisions
- jevos vs bart-large-mnli for zero-shot classification
- A yes/no LLM vs a fine-tuned BERT classifier
- jevos vs SetFit: zero-shot vs few-shot classification
- jevos vs Llama Guard for content safety checks
- jev serve vs llama.cpp server for classification
- jevos vs LM Studio: a decision server, not a chat app
- Local vs hosted LLM decisions: latency, cost, privacy
- A yes/no LLM vs a business rules engine
- LLM decisions vs keyword rules and regex
- The fastest AI model for yes/no decisions
- What makes a local LLM fast on a CPU
- Why one forward pass beats generating an answer
- Prefill vs decode: where LLM latency comes from
- Why LLM latency grows with the length of the text
- Why a hosted LLM API cannot answer in 50 ms
- Many questions about one text: why the extra ones are cheap
- CPU or GPU for a small LLM
- Latency budgets: where a 200 ms model fits
- Measuring LLM latency: median, p90 and warm-up
- Q4_K_M vs Q8_0: speed and size for a small model
- Throughput vs latency for a decision server
- What P(yes) means, and what it does not
- LLM calibration explained with yes/no answers
- Expected calibration error (ECE), explained
- Temperature scaling for LLM probabilities
- Platt scaling for a yes/no model
- Reading a reliability diagram
- How to choose a threshold for P(yes)
- Thresholds when a wrong yes costs more than a wrong no
- Human in the loop AI with a review band
- Precision and recall at a P(yes) threshold
- Base rates: why a 0.9 yes can still be wrong often
- Combining yes/no answers with AND, OR and NOT
- Logits, log-odds and P(yes)
- LLM confidence scores: probabilities vs self-reports
- How to write yes/no questions an LLM answers well
- Negation in yes/no questions for an LLM
- One condition per question: splitting compound questions
- Ask whether the text says it at all
- Scores as yes/no thresholds: is it at least high?
- Sending JSON as the text: designing the state
- Why wording changes an LLM's answer, and how to test it
- Mainly about: questions for messages with several topics
- Yes/no questions about tone and emotion
- Asking about intent: what does the writer want?
- Yes/no questions about long documents
- Using an English-only LLM with other languages
- Content moderation with a local LLM
- A Discord moderation bot with a local LLM
- Spam detection with yes/no questions
- Review moderation with a local LLM
- Email triage with a local LLM
- Support ticket routing with yes/no questions
- Urgency detection in customer messages
- Sentiment analysis with yes/no questions
- Intent detection with a local LLM
- Lead qualification with yes/no questions
- Fraud case triage with a local LLM
- Phishing email screening with a local LLM
- Log and alert triage with a local LLM
- Checking text for personal data with yes/no questions
- Prompt injection screening with a small model
- Document classification with a local LLM
- Product categorization with yes/no questions
- Contract clause detection with a local LLM
- Refund request triage with a local LLM
- Detecting cancellation intent in customer messages
- RAG evaluation with yes/no questions
- RAG faithfulness check with a local LLM
- Hallucination detection with a local LLM
- LLM regression tests in CI with yes/no checks
- Rubric design for an LLM judge
- Pairwise comparison with a yes/no judge
- LLM judge bias and how to control it
- Evaluation metrics for yes/no classifiers
- Building a yes/no test set for your own data
- Accuracy by kind of question: why one number hides failures
- Generating test questions with answers computed by code
- Benchmark contamination and truly held-out tests
- An LLM router with yes/no questions
- A model cascade: small model first, large model on doubt
- Semantic routing vs yes/no questions
- Gating AI agent tool calls with yes/no checks
- AI agent guardrails with yes/no questions
- Stop conditions for AI agents
- Logging LLM decisions for audit
- Reducing LLM cost with local yes/no decisions
- Replacing chat LLM calls with yes/no questions
- Structured output vs a probability
- A Python client for local LLM decisions
- Calling a local LLM decision server from JavaScript
- Local LLM yes/no decisions in n8n
- A Slack bot that uses local LLM decisions
- Home Assistant automations with local LLM decisions
- A LangChain tool for local yes/no decisions
- Batch decisions from files with jev decide
- Running LLM yes/no checks in GitHub Actions
- Securing a local LLM server with an API key
- curl examples for a local LLM decision API
- Self-hosted AI for decisions
- A private LLM for text classification
- On-premise LLM for business decisions
- GDPR and automated decision-making with an LLM
- Offline AI for decisions: no network needed
- Edge AI decisions on a CPU
- Run an LLM locally without a GPU
- Small language models explained
- When a small model is enough, and when it is not
- An LLM on a laptop: what it can do in real time
- What is GGUF, for someone deploying a classifier
- GGUF quantization types explained: Q4_K_M, Q8_0 and others
- GGUF vs safetensors
- llama.cpp vs Ollama for a classification service
- llama-cpp-python vs calling llama.cpp through ctypes
- llama.cpp on Windows without compiling
- Running llama.cpp CPU only
- Using llama.cpp prebuilt binaries instead of building