LLM Inference Capacity Planner
An interactive capacity-planning tool for local LLM inference. It estimates model weight storage, KV cache growth, runtime VRAM, minimum host memory, and GPU count from model size, quantization, context length, cache precision, batch size, and the target memory architecture. The breakdown makes the assumptions visible so a hardware plan can be adjusted before downloading a model or buying a GPU.
- Stack
- Next.js, TypeScript, LLM memory modeling
- Date
- 2026-10-04
- Source
- Interactive lab tool