← 05 / Writing
2026-09-21/local-ai9 min readLIVE

Prefill vs. Decode

Tokens per second is not one measurement. Prefill is the work of reading the prompt; decode is the work of generating output. Understanding both explains TTFT, output speed, bandwidth, compute, context length, and benchmark results.

When I ask whether a local LLM is fast, the first answer I usually get is a tokens-per-second number.

That number is incomplete.

An LLM request has two very different phases. First, the model reads the prompt. Then it generates the answer one token at a time.

Those phases are prefill and decode.

Prefill explains why a model can sit silently while processing a large prompt. Decode explains why the answer then appears at 10, 40, or 80 tokens per second. The two phases stress hardware differently, so one tokens-per-second number cannot describe both.

The short version

TermMeaning
PrefillReading and processing the input prompt
DecodeGenerating output tokens one at a time
Prompt tokens/sPrefill throughput
Output tokens/sDecode speed
TTFTTime to first token
TPOT / ITLTime between generated tokens
TFLOPSPeak mathematical throughput
GB/s or TB/sMemory bandwidth

The important distinction is simple:

  • Prefill tends to benefit from compute throughput and efficient matrix/attention kernels.
  • Decode at batch size one is often limited by memory bandwidth.
  • Longer occupied context makes prefill take longer and can also slow decode.
  • Tokens/s is ambiguous unless the benchmark says whether it means prompt processing, single-stream output, or aggregate server throughput.

One request, two different jobs

A basic request looks like this:

Prompt → Prefill → First token → Decode → Stream token → Decode → ...

If I send a 40,000-token repository snapshot and ask for a 500-token answer, the model does not treat all 40,500 tokens the same way.

During prefill, it can process many prompt tokens in parallel. During decode, each new output token depends on the token immediately before it. The model must finish step 1 before step 2, then finish step 2 before step 3.

That sequential dependency is why decode is harder to parallelize for one conversation.

Prefill: reading the prompt

Prefill is the first forward pass over the input. The model processes the prompt and builds the attention state that later decode steps reuse.

The key feature is parallelism. The entire prompt already exists, so many input tokens can be processed together as large matrix operations.

Prefill throughput is commonly reported as:

prompt tokens/s = input tokens / prefill time

If a 50,000-token prompt takes 25 seconds to process, the prefill rate is about 2,000 tokens/s. The user still waits about 25 seconds before useful output starts.

That wait contributes heavily to time to first token, or TTFT:

TTFT ≈ queue time + tokenization + prefill + first decode step

Large matrix operations give the processor enough parallel work to use its compute units efficiently. That is why prefill is usually associated with matrix throughput, attention kernels, kernel fusion, and low-precision compute.

Decode: writing the answer

Decode begins after prefill has processed the prompt.

At each step, the model accepts the newest token, runs it through the network, produces probabilities for the next token, selects one, and repeats.

Unlike prefill, each output token depends on the previous one. For one conversation, that makes decode much harder to parallelize.

At batch size one, the model often needs to stream most of its weights from memory again for every generated token. That is why memory bandwidth becomes such an important predictor.

Output speed is:

output tokens/s = generated tokens / decode time

The inverse is time per output token, or TPOT:

TPOT ≈ 1 / output tokens/s

For example, 40 tokens/s is about 25 ms per token. Ten tokens/s is about 100 ms per token.

Inter-token latency, or ITL, describes the spacing between streamed tokens and is closely related to TPOT.

Why memory bandwidth predicts decode

Memory bandwidth is the rate at which the GPU can move data from its main memory into the compute cores. It is normally listed in GB/s or TB/s.

A useful first-order estimate is:

ideal decode tokens/s ≈ memory bandwidth / model weight size

This assumes the hardware reads the model once per token, reaches 100 percent of advertised bandwidth, and pays no cost for the KV cache, activations, quantization, sampling, or kernel launches. None of those assumptions is fully true.

It is still a valuable ceiling.

Suppose a quantized model occupies about 18 GB. A very rough decode ceiling is:

ideal decode tokens/s ≈ memory bandwidth / model weight size

A system with 600 GB/s of effective memory bandwidth has a much lower theoretical ceiling than one with 1.2 TB/s, assuming the same model and backend.

This is not a benchmark. Real speed is lower because the runtime also moves KV-cache data and activations, performs sampling, launches kernels, and may need to unpack quantized weights.

The formula is still useful because it explains why memory bandwidth often predicts batch-one decode better than peak TFLOPS.

Why CPU offload hurts

If the whole model does not fit in accelerator memory, an inference engine can offload some layers to system RAM or the CPU.

That solves capacity, but usually hurts decode speed because every generated token must repeatedly cross the slower part of the memory path.

TFLOPS: the most misunderstood number

A FLOP is one floating-point operation. A TFLOP is one trillion floating-point operations per second.

Peak TFLOPS is only meaningful when I also know the precision being measured—FP32, FP16, BF16, FP8, FP4, and so on—and whether the workload can actually use the hardware path behind that number.

A device can advertise enormous low-precision throughput and still perform poorly if the model format or inference backend cannot use the right kernels.

TFLOPS matters most when the workload is compute-bound, which is commonly the case during prefill. It is a much weaker standalone predictor of batch-one decode.

The roofline model in one formula

The relationship between compute and bandwidth can be summarized with the roofline model:

attainable compute ≤ min(
    peak compute,
    memory bandwidth × arithmetic intensity
)

Arithmetic intensity means how much math is performed for each byte moved from memory.

Prefill can reuse data across many input tokens, so arithmetic intensity is relatively high and the workload can move toward the compute limit.

Batch-one decode has much less weight reuse, so it often hits the memory-bandwidth limit first.

Batching several sequences increases reuse and can move decode back toward a more compute-heavy workload.

Context length changes both phases

Context length is not only a memory setting.

During prefill, every occupied input token must be processed. Longer prompts therefore increase the work required before the first token appears.

During decode, the model repeatedly attends to cached history. A longer occupied context means more KV-cache data may need to be read, so output speed can also fall as the conversation grows.

The important distinction is:

configured context = maximum context the runtime allows
occupied context   = tokens actually present in the request

Starting a server with a 262K context limit and sending a 500-token prompt is not a 262K-context benchmark.

KV cache and prefix caching

The KV cache stores attention state created during prefill so the model does not recompute the entire prompt for every output token.

Its memory use grows roughly with:

tokens × attention layers × KV heads × head size
       × K and V × bytes per value × parallel sequences

Grouped-query attention, sliding-window attention, hybrid architectures, and lower-precision KV caches can reduce that cost.

Prefix caching reuses previously computed state when the beginning of a prompt stays identical. That can dramatically reduce repeated prefill work in chats and tool-using workflows.

The important limitation is that the prefix must remain stable. Changing an early system message or tool definition can invalidate everything after that point.

Quantization and performance

Quantization reduces the number of bits used to store model weights. Smaller weights mean fewer bytes must be moved during decode, so quantization can improve both memory usage and output speed.

The speedup is not perfectly proportional to model size because quantized kernels must unpack and rescale values, and different backends optimize different formats differently.

That is enough for this article; quantization deserves its own deeper explanation.

Batching changes the answer

Batching lets the model process several sequences together.

During decode, the same loaded weights can serve multiple sequences. This increases weight reuse and raises total throughput, but it can also increase per-request latency.

That means a server benchmark can look much faster than a single-user experience.

When reading a benchmark, check whether tokens/s is:

  • per request, or
  • aggregate across many requests.

A server producing 1,000 total tokens/s across 100 users is not giving each user a 1,000-token/s conversation.

Speculative decoding and MTP

Normal autoregressive decode accepts one token per full model step.

Speculative decoding and multi-token prediction (MTP) try to produce several useful tokens from less target-model work. A draft mechanism proposes future tokens and the main model verifies them.

This can improve decode throughput, but the gain depends on proposal accuracy, verification efficiency, and extra memory overhead.

These techniques do not increase physical memory bandwidth. They improve how much useful output is extracted from each expensive pass.

How to read an LLM benchmark

A single tokens/s number is not enough.

A useful benchmark should tell me:

MeasurementWhy it matters
Model and precisionDetermines weight size and kernel path
Occupied prompt lengthDetermines how much prefill work was done
Prompt tokens/sMeasures prefill throughput
TTFTMeasures how long the user waits for output to begin
Output tokens/sMeasures decode speed
TPOT / ITLMeasures streaming latency
Batch size / concurrencyDistinguishes single-user latency from server throughput
KV-cache typeAffects memory use and bandwidth
Backend and versionDifferent runtimes use different kernels
OffloadShows whether the whole model stayed on the accelerator
Speculative settingsReveals whether draft decoding or MTP changed the result

For meaningful comparisons, use at least two prompt lengths: one representative everyday prompt and one long-context case.

The practical takeaway

Prefill and decode are different workloads.

Prefill is mostly about processing many input tokens efficiently. It benefits from parallel matrix compute, optimized attention, and good kernels.

Decode is sequential. At batch size one, it is often limited by how quickly the system can stream model weights and KV-cache data through memory.

That is why one tokens-per-second number can be misleading.

When evaluating local LLM performance, I want to know four things:

  1. How long does prefill take at the prompt sizes I actually use?
  2. What is the time to first token?
  3. How fast does decode run for one interactive request?
  4. How do those numbers change as occupied context grows?

Once those are separated, benchmark results become much easier to understand.

Sources