LLM Quantization Explained
FP16, BF16, FP8, Q8, Q4, GGUF, KV cache, prefill, and decode all describe different parts of local AI. Here is how they fit together—and what they mean for memory, speed, and model quality.
The local-AI world has a vocabulary problem.
A model page says BF16. A download says Q4_K_M. A benchmark says W4A16. Someone calls FP8 “unquantized,” another person calls Q8 “basically lossless,” and the model still runs out of memory even though its file fits on the GPU.
None of those statements is completely absurd. They are just describing different parts of the system.
The most important distinction is this: precision describes how numbers are represented; quantization is a process for mapping numbers into a smaller representation. The number of bits affects model size, but it does not tell me everything about quality, speed, or memory use.
This is the guide I wish model download pages included.
The short version
| Term | What it actually means | Practical effect |
|---|---|---|
| FP32 | 32-bit floating-point numbers | High numerical precision; usually wasteful for LLM inference |
| FP16 | 16-bit floating point | Half the weight memory of FP32; common GPU format |
| BF16 | 16-bit floating point with FP32-like range | Common training and source-checkpoint format |
| FP8 | An 8-bit floating-point format | About one byte per raw value, but needs hardware and software support |
| INT8 / Q8 | An 8-bit integer-style quantization | Similar nominal bit count to FP8, but a different representation and execution path |
| Q6 / Q5 / Q4 / Q3 | Rough quantization tier | Fewer bits reduce weight memory and usually increase quantization error |
| BPW | Average bits per weight | More honest than the digit in a quant name when formats use mixed types and metadata |
| GGUF | A model container used heavily by llama.cpp | Can package tensors, tokenizer data, metadata, and quantized weights |
| KV cache | Stored attention state for previous tokens | Grows with context length and can consume the memory left after loading weights |
| Prefill | Processing the prompt or existing context | Determines how quickly generation can begin |
| Decode | Generating new tokens one at a time | The speed I feel while reading the answer |
If I remember only four things, they are these:
Q4is not the same thing asFP4, andQ8is not the same thing asFP8.- A 4-bit model does not necessarily use exactly 4 bits for every stored value.
- Quantizing weights does not automatically quantize the KV cache, activations, or computation to the same bit width.
- The largest model file that fits is not necessarily a usable configuration.
Start with weights and parameters
An LLM is mostly a very large collection of learned numbers called parameters or weights. During training, the model adjusts those numbers until it becomes good at predicting the next token.
When a model is described as 7B, 27B, or 70B, the B means billions of parameters. It does not mean gigabytes.
The basic weight-size estimate is:
weight storage = parameter count × bits per weight ÷ 8
A dense 27-billion-parameter model therefore has these theoretical storage floors:
| Weight representation | Nominal bits per weight | Theoretical size for 27B parameters |
|---|---|---|
| FP32 | 32 | 108 GB |
| FP16 / BF16 | 16 | 54 GB |
| FP8 / INT8 / Q8 | 8 | 27 GB |
| 6-bit quant | 6 | 20.25 GB |
| 5-bit quant | 5 | 16.88 GB |
| 4-bit quant | 4 | 13.5 GB |
| 3-bit quant | 3 | 10.13 GB |
| 2-bit quant | 2 | 6.75 GB |
Those are clean mathematical estimates in decimal gigabytes. Real files are different because they can include higher-precision tensors, scales, zero points, block metadata, tokenizer data, alignment, and other overhead. A format advertised as “4-bit” may average 4.5 or 5 bits per weight. The actual file size is more useful than the label.
It is also worth keeping GB and GiB straight. Drive makers normally use decimal gigabytes. Operating systems and memory tools may report binary gibibytes. One GiB is 1,073,741,824 bytes, so the same file can appear to have two different sizes without changing at all.
What precision means
Precision is how much information a numeric format can preserve. A number stored with more bits can generally represent either more distinct values, a wider range, or some combination of both.
Floating-point numbers divide their bits among three jobs:
- A sign says whether the value is positive or negative.
- An exponent controls the range—how extremely large or small the value can be.
- A mantissa, also called a significand or fraction, controls how finely values can be distinguished.
This is why two 16-bit formats can behave differently:
| Format | Total bits | Main tradeoff |
|---|---|---|
| FP32 | 32 | Wide range and high precision |
| FP16 | 16 | More fraction precision than BF16, but a much smaller numeric range |
| BF16 | 16 | FP32-like range with fewer fraction bits |
| FP8 E4M3 | 8 | More precision, less range than E5M2 |
| FP8 E5M2 | 8 | More range, less precision than E4M3 |
BF16 is popular for training because it keeps the eight exponent bits of FP32 while using fewer fraction bits. PyTorch describes it as having FP32's dynamic range with the same two-byte footprint as FP16. That does not make BF16 universally “better” than FP16; it means the two formats spend their 16 bits differently.
The phrase full precision is dangerously vague. In traditional numerical computing it often means FP32. In modern LLM discussions it frequently means the original BF16 or FP16 checkpoint—the model before low-bit post-training quantization. I always check which meaning the author intended.
What quantization actually does
Imagine a group of weights ranging from -1.0 to 1.0. In FP16 or BF16, the model can represent many values inside that interval. A simple 4-bit integer quantizer has only 16 codes available.
The quantizer chooses a scale, maps each original value to the closest code, and stores that smaller code. During inference, the runtime uses the scale to reconstruct an approximation of the original value or dequantizes blocks as they enter a matrix multiplication.
In simplified form:
quantized = round(original / scale) + zero_point
approximate_original = (quantized - zero_point) × scale
The difference between the original value and its reconstruction is quantization error.
Good quantization is much more sophisticated than applying one scale to an entire model. Modern methods may use:
- A separate scale for every tensor, channel, row, or small block
- Symmetric ranges around zero or asymmetric ranges with a zero point
- Calibration data to find which weights are most sensitive
- Mixed bit widths that protect important tensors
- Non-uniform codebooks that devote more values to common regions
- Higher-precision accumulation even when stored weights are low-bit
This is why two files with “4-bit” in their names can have noticeably different sizes and quality.
Stored precision is not compute precision
This is the distinction that clears up most confusion.
An inference operation involves several kinds of numbers:
| Component | What it is | Possible format |
|---|---|---|
| Weights | The learned model parameters | Q4, INT8, FP8, FP16, BF16 |
| Activations | Intermediate values created while the model runs | FP16, BF16, FP8, INT8 |
| Accumulators | Running totals inside matrix multiplication | FP16 or FP32, often higher precision than inputs |
| KV cache | Attention state retained for prior tokens | FP16, BF16, Q8, Q4, or another cache-specific format |
W4A16 means 4-bit weights and 16-bit activations. W8A8 means 8-bit weights and 8-bit activations. Neither label necessarily tells me the accumulator format.
A GGUF Q4 model is commonly stored in compressed blocks, then processed by kernels designed for those blocks. Parts may be dequantized to a higher precision during computation. The whole model does not suddenly become a pure 4-bit computer from disk to final token.
The same warning applies to FP8. A model described as FP8 may store weights in FP8 while using BF16, FP16, or FP32 for some activations, sensitive operations, and accumulations. The exact path depends on the model, runtime, kernel, and hardware.
FP8 is not Q8, and FP4 is not Q4
The digit tells me the storage width. The letters tell me how those bits are interpreted.
FP8 stores a sign, exponent, and fraction in eight bits. Because it has an exponent, it can cover values across a wide range, but with very few steps at any one magnitude.
INT8 or Q8 usually stores integer codes plus scales for blocks or channels. It can offer fine resolution within a chosen local range, but the scale metadata and grouping strategy matter.
Q4 in a GGUF filename normally refers to a block-quantization family used by the GGML/llama.cpp ecosystem. It does not mean IEEE-style 4-bit floating point.
Two representations can both average eight bits per weight and still differ in:
- Which values they represent well
- How they handle outliers
- Whether activations are also quantized
- What hardware instructions they can use
- How much metadata they need
- Which runtime can execute them efficiently
Bit count is one dimension, not an identity.
How to read a GGUF quant name
GGUF is a container format, not a single quantization algorithm. It can hold model tensors in different data types along with architecture metadata and tokenizer information.
Names vary as the ecosystem evolves, but these are common patterns:
| Example | Practical interpretation |
|---|---|
F16 | Weights stored primarily as 16-bit floating point |
Q8_0 | An 8-bit block quant; large and usually close to the source model |
Q6_K | A K-quant around the 6-bit tier |
Q5_K_M | A 5-bit-class K-quant using a mixed, larger preset |
Q4_K_S | A smaller Q4 K-quant preset |
Q4_K_M | A larger Q4 K-quant preset that protects more tensors |
Q3_K_L | A larger preset in the 3-bit K-quant family |
IQ3_* / IQ4_* | Importance-aware quantization families designed for strong quality at low average BPW |
S, M, and L generally mean small, medium, and large variants within a family. They do not create a universal quality scale across every model and quantizer.
The K in K-quants refers to a family of block formats that can use different quant types for different tensors. That is one reason Q4_K_M does not land at exactly 4.00 bits per weight.
An importance matrix, often shortened to imatrix, measures how much different weights or tensor regions matter on calibration data. The quantizer can then preserve sensitive parts more carefully. The llama.cpp tooling supports creating and applying importance matrices during quantization.
My download order is:
- Confirm the model and architecture.
- Check the actual file size and expected runtime memory.
- Choose a well-tested quant family supported by my runtime.
- Prefer a quant made with representative calibration data.
- Compare task benchmarks or run my own tests.
The filename is a clue, not a quality guarantee.
The major quantization approaches
These names are often mixed together even though some are algorithms, some are data types, and some are file formats.
| Name | What it is | Where I commonly see it |
|---|---|---|
| PTQ | Post-training quantization performed after the model is trained | Most downloadable local-model quants |
| QAT | Quantization-aware training that simulates low precision during training | Models intentionally trained to tolerate a target precision |
| GPTQ | A post-training, usually weight-only quantization method using approximate second-order information | CUDA-focused inference stacks |
| AWQ | Activation-aware weight quantization that protects salient weights | GPU serving engines and Transformers ecosystems |
| NF4 | A 4-bit NormalFloat data type designed around normally distributed neural-network weights | bitsandbytes, especially QLoRA |
| GGUF | A container that can store many tensor and quant types | llama.cpp, LM Studio, compatible local runtimes |
| Safetensors | A safe tensor storage format, not a quantization level | Original and quantized Transformers checkpoints |
Post-training quantization is the normal path for local inference: start from a trained BF16 or FP16 model and compress it without retraining the whole network.
Quantization-aware training exposes the model to simulated low-precision behavior during training or fine-tuning. It can recover quality, but it is more expensive than converting an existing checkpoint.
GPTQ and AWQ are methods for deciding how weights should be quantized. NF4 is a data type introduced with QLoRA. GGUF and Safetensors are containers. Comparing “GGUF versus GPTQ” is partly a comparison of ecosystem and packaging, not simply one quantization equation against another.
Does lower precision make the model worse?
Usually, but not in a clean straight line.
Quantization changes the model's weights, so some loss is inevitable unless the representation happens to preserve every value that matters. The practical question is whether that loss changes the work I care about.
At 8 bits, a good quant is often difficult to distinguish from the source model in ordinary use. At 6 and 5 bits, quality is commonly very close. Four bits is the local-inference sweet spot because it cuts weight memory dramatically while retaining much of the model's capability. Three bits can be impressive, but mistakes and instability become more model- and task-dependent. Two-bit and ternary formats are far more sensitive to the architecture and quantization recipe.
That is a rule of thumb, not a law.
A larger model at Q4 can outperform a smaller model at FP16 because parameter count, architecture, and training quality may matter more than the precision difference. A strong 27B Q4 is not automatically worse than a 14B BF16. But a Q4 version of the same 27B model should be expected to lose something relative to its own BF16 source.
The loss may show up as:
- Slightly worse perplexity
- More formatting or repetition problems
- Reduced factual recall
- Less reliable coding or mathematics
- Greater fragility on long prompts
- A higher chance of choosing the wrong tool or argument in agent workflows
Average benchmark scores can hide rare but important failures. For coding, I care about whether the quant can complete a multi-step edit, recover from a failed test, and preserve exact identifiers—not only whether it scores within one point on a broad benchmark.
Perplexity is useful, but it is not the whole answer
Perplexity measures how surprised a model is by the next tokens in evaluation text. Lower is better. Quantization papers often compare perplexity before and after compression because it is repeatable and sensitive to degradation.
It is not a complete measure of usefulness. A small perplexity change may or may not affect code generation, tool calling, instruction following, or creative writing. Benchmark accuracy may also fail to capture how often an agent gets stuck after its first mistake.
My preferred test set contains real prompts from my workflow:
- A focused bug fix
- A multi-file refactor
- A repository question with irrelevant files nearby
- Structured JSON or tool calls
- A long conversation near my normal context size
- A task that requires reading a test failure and trying again
The best quant is the smallest one that still passes those tests reliably.
Quantization can improve speed—but not automatically
LLM decoding frequently moves a large portion of the model's weights through memory for every generated token. If I reduce the weight size, the hardware has fewer bytes to read. That is why memory bandwidth is such an important predictor of single-user token generation.
A rough intuition is:
maximum decode rate ≈ memory bandwidth ÷ effective model bytes read per token
Real performance is lower because kernels, caches, synchronization, architecture, and compute also matter.
Lower-bit weights can be faster when:
- The workload is limited by memory bandwidth
- The runtime has optimized kernels for that exact quant format
- The entire model stays in fast memory
- Dequantization overhead is small
They can be slower when:
- The hardware lacks native or optimized low-bit support
- Conversion and dequantization cost more than the memory savings
- Some layers spill into system RAM
- The model is split across a slow interconnect
- The prompt-processing phase is compute-bound instead of bandwidth-bound
Never assume Q3 must be faster than Q4, or FP8 must be faster than BF16, on every device. The supported kernel matters as much as the nominal bit width.
The model file is only the first memory bill
Loading weights is not the same as running a useful model. Total inference memory includes:
total memory ≈ weights + KV cache + runtime buffers + temporary activations + system overhead
KV cache
Autoregressive models generate one token at a time. Without a cache, the model would repeatedly recompute attention information for all previous tokens. The key-value cache stores that state so it can be reused.
The cache grows with the number of tokens and usually with the number of concurrent sequences. Its size depends on the architecture, layer count, KV heads, head dimension, cache data type, and context length. Grouped-query attention, multi-query attention, sliding-window attention, and hybrid architectures can change the memory curve substantially.
Quantizing model weights does not automatically shrink this cache. A Q4 model may still use an FP16 KV cache unless I configure a supported cache quantization mode.
Runtime buffers and activations
The inference engine needs working memory for computation graphs, temporary tensors, attention operations, token batches, and backend-specific buffers. Multimodal models may also load a vision encoder or projector.
Context allocation
Some runtimes allocate KV capacity for the configured context window up front. Asking for 256K context can therefore consume memory even when the current prompt is short. Other runtimes grow the cache dynamically.
This is why a 13.5 GB theoretical Q4 weight floor does not make a 27B model comfortable on a 16 GB card. It may load, but there still has to be room for the conversation.
Context, tokens, and the terms around them
A token is a unit produced by the model's tokenizer. It may be a whole word, part of a word, punctuation, whitespace, or a fragment of source code. Token count does not map perfectly to word count, and code can tokenize very differently from prose.
The context window is the maximum token span the runtime and model can consider for the current request. It includes more than the latest user message:
- System instructions
- Chat history
- Retrieved documents
- Source files
- Tool definitions and tool results
- The model's generated reasoning or answer, depending on the system
An advertised 128K context window does not mean all 128K tokens are equally useful. Models can lose retrieval accuracy or instruction focus long before reaching the hard limit. Usable context is a quality question as well as a memory question.
RoPE scaling and similar techniques attempt to extend a model beyond its trained positional range. They can make a longer window technically possible, but they do not guarantee that the model reasons equally well across it.
Prefill, decode, latency, and throughput
Local benchmarks become much easier to read once the generation process is split in two:
flowchart LR
P["Prompt and context"] --> T["Tokenization"]
T --> F["Prefill: process input tokens"]
F --> K["KV cache"]
K --> D["Decode: generate one token"]
D --> K
D --> O["Stream output"]
Prefill, also called prompt evaluation, processes the input tokens and builds the KV cache. It is highly parallel and often compute-heavy. Prompt-evaluation speed is usually reported in tokens per second.
Decode, also called generation or token evaluation, produces new tokens autoregressively. For a single conversation, it is often limited by memory bandwidth because the model weights must be consulted for every token. This is the 13 tok/s or 60 tok/s number that determines how fast text appears.
Time to first token (TTFT) combines queueing, tokenization, model loading if necessary, and prefill before the first output token appears.
Inter-token latency (ITL) measures the delay between generated tokens. It is the inverse of per-user decode speed.
Throughput is total work completed across users or requests. Batching may improve total throughput while making an individual request wait longer.
Batch size is how many sequences or token groups are processed together. Concurrency is how many requests are active. Parallel slots in a local server reserve capacity for multiple conversations and can multiply KV-cache requirements.
Dense models, MoE models, and active parameters
A dense model uses essentially all of its main parameters for every token. A 27B dense model therefore consults roughly the full 27B-parameter network during generation.
A mixture-of-experts (MoE) model contains many expert blocks but routes each token through only a subset. A model might have 100B total parameters and 10B active parameters per token.
This creates two separate numbers:
- Total parameters strongly influence storage and memory capacity.
- Active parameters strongly influence compute per token.
MoE does not mean the inactive weights disappear. Unless the runtime streams experts from slower storage, the full model still needs to reside somewhere. “100B total, 10B active” can compute more like a smaller model while still demanding memory like a 100B model.
GPU offload, CPU offload, and unified memory
GPU offload moves model layers or operations from the CPU to the GPU. In llama.cpp, offloading more layers normally improves speed until VRAM is full.
CPU offload keeps some weights in system RAM and executes or feeds those layers through the CPU path. It can make a model run that would not fit entirely in VRAM, but crossing PCIe or computing on the CPU can sharply reduce speed.
Unified memory systems let the CPU and GPU access one physical memory pool. That can make very large models easier to load without explicit PCIe copies, but capacity and bandwidth are still separate. A machine with 128 GB of unified memory is not equivalent to a GPU with 128 GB of high-bandwidth VRAM.
Tensor parallelism splits tensor operations across accelerators. Pipeline parallelism assigns different layer ranges to different devices. Both can make large models possible, but communication speed between devices becomes part of every token.
Inference, training, fine-tuning, LoRA, and QLoRA
These terms describe different workloads with very different memory requirements.
Inference means running a trained model to generate an answer. It needs weights, a KV cache, and runtime working memory.
Training updates all or most model weights. It also needs gradients, saved activations, optimizer states, and often higher-precision master weights. Training can require many times more memory than inference.
Fine-tuning continues training an existing model on a new dataset or behavior target.
LoRA keeps the original model mostly frozen and trains small low-rank adapter matrices. It dramatically reduces the number of trainable parameters.
QLoRA loads the frozen base model in a 4-bit format—commonly NF4—and trains LoRA adapters through that quantized base. It reduces fine-tuning memory, but “I can run this model in Q4” still does not mean “I can fully train this model in the same memory.”
Base, instruct, reasoning, and agentic models
Base model means the pretrained next-token predictor before chat or instruction alignment. It may be excellent for further training but awkward as an assistant.
Instruct or chat model means the base has been tuned to follow requests and conversation roles. It usually expects a specific chat template.
Reasoning model is trained or prompted to spend more tokens solving multi-step problems. More generated reasoning can improve difficult tasks, but it also increases latency and context use.
Agentic model is not a numeric format. It describes a model or training recipe aimed at planning, calling tools, editing files, reading results, and continuing across multiple steps. The surrounding harness is just as important as the weights.
Chat template is the exact formatting that converts roles such as system, user, assistant, and tool into the token sequence the model was trained to expect. A wrong template can make a good model look broken.
Temperature, top-p, and sampling
After the model produces a score for every possible next token, the runtime must choose one.
Logits are the raw scores. Softmax converts them into probabilities.
Temperature reshapes that distribution. Lower values make high-probability tokens more dominant; higher values make unlikely choices more competitive. Temperature does not add knowledge or intelligence.
Top-p, or nucleus sampling, restricts sampling to the smallest set of tokens whose combined probability reaches a threshold. Top-k restricts the choice to the k highest-scoring tokens.
Greedy decoding always chooses the most likely next token. It is deterministic for the same inputs and runtime conditions, but it is not automatically the most accurate strategy for every task.
A seed initializes the pseudo-random sampler. The same seed improves reproducibility, but different hardware, kernels, batching, or floating-point behavior can still cause outputs to diverge.
My practical quantization ladder
If I am choosing a local model and have not benchmarked it yet, this is where I start:
| Goal | Starting point |
|---|---|
| Maximum fidelity or a reference baseline | BF16 or FP16 |
| Near-source inference with plenty of memory | Q8 or a well-supported FP8 checkpoint |
| High quality with meaningful savings | Q6 or Q5 |
| Best general local-inference tradeoff | A strong Q4, often Q4_K_M or a good importance-aware equivalent |
| Fit a model one tier above the hardware | A carefully tested Q3 |
| Experimental maximum compression | Q2, ternary, or architecture-specific ultra-low-bit formats |
Then I verify four things:
- Fit: weights, KV cache, and runtime buffers all fit at my real context size.
- Support: my runtime has optimized kernels for that exact format and hardware.
- Speed: I measure both prompt evaluation and token generation.
- Quality: I test the model on my own coding, tool-use, or writing workflow.
I would rather run a stable Q4 with 32K of useful context than a Q6 that leaves the machine one long prompt away from an out-of-memory error. I would also rather use a slightly smaller model that reliably calls tools than a giant quant that becomes confused halfway through an agent loop.
Precision is a resource. Quantization decides how aggressively I spend it.
Sources
- PyTorch: What every user should know about mixed-precision training
- PyTorch: BF16 on Intel Xeon processors
- Hugging Face Transformers: How caching works
- Hugging Face Transformers: KV-cache strategies
- Hugging Face Transformers: Bitsandbytes quantization
- llama.cpp: Importance-matrix quantization tooling
- GGUF specification
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- QLoRA: Efficient Finetuning of Quantized LLMs