← 02 / Lab records
LAB-001/2026-10-04ACTIVE

LLM Inference Capacity Planner

Estimate the memory footprint of an LLM inference workload before you choose hardware. Configure the model, cache, and target system; the report updates immediately.

Configuration

Workload profile

Live estimate
00 / Model source

Hugging Face model

Model search?
HF Hub
01 / Runtime memory

Context & cache

The controls that most directly change inference memory after choosing a model and quantization.

02 / Hardware target

Memory architecture

System type?
GPU model?
52 options
03 / Model metadata

Imported data & overrides

Populated from Hugging Face when available. Keep these controls for repositories with incomplete metadata or manual estimates.

Weight quantization?
Capacity report
49.6 GiBMinimal VRAM required
324 GiB GPUs
Model weights
36.7 GiB
Quantized parameters in memory
KV cache
10.7 GiB
32,768 tokens × 1 sequence × 84 attention layers
Compute buffers
2.20 GiB
Allocator and inference workspace
Runtime subtotal
49.6 GiB
Before 10% safety margin
Recommended VRAM
54.5 GiB
Minimum required + 10% safety margin
Minimum host RAM47.4 GiB
On-disk model37.8 GiB
Run this configuration

llama.cpp server command

Uses your context, K/V cache precision, and concurrent sequences from above. Install a current llama.cpp build with the GPU backend for your hardware.

Text-only mode skips the vision projector to save memory. Image input is disabled.

More launch settings

Prompt batch and micro-batch sizes control prompt processing, separate from parallel conversations. Smaller micro-batches can reduce workspace memory.

32,768 tokens/slot × 1 slot
llama-server \
  -m '/path/to/model.gguf' \
  -c 32768 \
  -np 1 \
  --cache-type-k f16 \
  --cache-type-v f16 \
  -ngl all \
  --flash-attn on \
  --batch-size 512 \
  --ubatch-size 128 \
  --split-mode layer \
  --no-mmproj \
  --host 127.0.0.1 \
  --port 8080

Total context (-c) is 32,768 tokens; -np sets parallel slots. All model layers are requested on GPU. Layer splitting requires enough visible GPUs to meet the report. Cache quantization and Flash Attention support depend on the model and GPU backend; 16/8/4-bit selections use f16/q8_0/q4_0 respectively, and quantized cache storage includes block metadata. Open http://127.0.0.1:8080 after startup.

llama.cpp server flag reference
MODEL ASSUMPTION // Approx. 84 layers × 8,333 hidden size. Real usage varies by architecture, inference engine, tensor parallelism, and quantization format. Treat this as a planning estimate, not a compatibility guarantee.