LLM Inference Capacity Planner
Estimate the memory footprint of an LLM inference workload before you choose hardware. Configure the model, cache, and target system; the report updates immediately.
Workload profile
Hugging Face model
Context & cache
The controls that most directly change inference memory after choosing a model and quantization.
Memory architecture
Imported data & overrides
Populated from Hugging Face when available. Keep these controls for repositories with incomplete metadata or manual estimates.
- Model weights
- 36.7 GiB
- Quantized parameters in memory
- KV cache
- 10.7 GiB
- 32,768 tokens × 1 sequence × 84 attention layers
- Compute buffers
- 2.20 GiB
- Allocator and inference workspace
- Runtime subtotal
- 49.6 GiB
- Before 10% safety margin
- Recommended VRAM
- 54.5 GiB
- Minimum required + 10% safety margin
llama.cpp server command
Uses your context, K/V cache precision, and concurrent sequences from above. Install a current llama.cpp build with the GPU backend for your hardware.
Text-only mode skips the vision projector to save memory. Image input is disabled.
More launch settings
Prompt batch and micro-batch sizes control prompt processing, separate from parallel conversations. Smaller micro-batches can reduce workspace memory.
llama-server \
-m '/path/to/model.gguf' \
-c 32768 \
-np 1 \
--cache-type-k f16 \
--cache-type-v f16 \
-ngl all \
--flash-attn on \
--batch-size 512 \
--ubatch-size 128 \
--split-mode layer \
--no-mmproj \
--host 127.0.0.1 \
--port 8080Total context (-c) is 32,768 tokens; -np sets parallel slots. All model layers are requested on GPU. Layer splitting requires enough visible GPUs to meet the report. Cache quantization and Flash Attention support depend on the model and GPU backend; 16/8/4-bit selections use f16/q8_0/q4_0 respectively, and quantized cache storage includes block metadata. Open http://127.0.0.1:8080 after startup.
llama.cpp server flag reference