← 05 / Writing
2026-09-20//llm /local-ai9 min readLIVE

Local Coding Models by Memory

From 4 GB to 512 GB, the right local coding model is the one that leaves enough memory for context, KV cache, and the runtime—not the largest file that barely loads.

Most local LLM guides answer the easiest question: what is the largest model file that fits in memory?

That is not the question I care about.

I want to know which model can read the code, hold a useful conversation, run tools, survive a few failed tests, and still have enough memory left to finish the task. A model that consumes the entire card before the first prompt is not a working setup. It is a successful loading screen.

I built this list from Cloud Codes' video, Best Local Coding AI for Your GPU (4GB to 512GB), then checked the linked model cards and current quantization sizes. The exact winner will keep changing, but the memory math will not.

The short version

These are the configurations I would start with today. Weight sizes are approximate and do not include the KV cache or runtime overhead.

MemoryDaily-driver pickQuant / weight targetWhat it is good for
4 GBSpark-X2.5-4B4-bit, about 2.6 GBExplaining code, one focused bug, small edits
6 GBNeoHorse-1-4BQ5_K_M, 3.07 GBFocused coding with better tool use
8 GBOrnith-1.5-9B4-bit, about 5.75 GBOne file plus a useful conversation
12 GBOrnith-1.5-9B8-bit, about 9.75 GBThe same model with cleaner output; still keep context lean
16 GBQwen3.8-27B GSQ-RCOIQ3_S, 11.8 GBReal coding, limited multi-file work, 16K–32K context
24 GBQwen3.8-27BGood 4-bit / low-5-bit build, about 16.5 GBMulti-file changes, tests, longer conversations
32 GBQwen3.8-27B6-bit, about 22 GBHigh-quality 27B with comfortable working room
48 GBQwen3-Coder-NextCompact IQ4/dynamic quant, about 38 GBLong agent loops, editing, testing, recovery
64 GBQwen3-Coder-NextOfficial Q4_K_M, 48.4 GBThe practical autonomous-agent tier
80 GBQwen3-Coder-NextQ6_K, 65.5 GBBetter fidelity with room for a large context
96 GBQwen3-Coder-NextQ8_0, 84.8 GBNear-lossless coding agent on one large card
128 GBQwen3.8-Flash-NextQ4-class, about 111 GBLarge sparse model with enough room to work
192 GBGLM-5.3-FlashLower 4-bit build, roughly 162 GBFrontier-class coding without filling the whole machine
256 GBMiniMax-M3IQ4_XS, about 208 GB428B total / 23B active; strong all-day sparse-agent setup
512 GBGLM-5.33-bit, about 343 GBThe sensible daily driver; Q4 at 467 GB is the maximum-quality option

There are two places where I would resist the obvious upgrade.

At 12 GB, I would rather run a good 9B model at 8-bit than squeeze in a larger model with almost no context. At 512 GB, I would rather run GLM-5.3 at a good 3-bit quant with more than 150 GB free than load the 467 GB Q4 build and pretend the remaining memory is generous.

The memory nobody puts in the model name

Three things compete for the same memory:

  1. Weights — the GGUF or other model files.
  2. KV cache — the model's working notes for every token already in the conversation.
  3. Runtime scratch space — temporary buffers, graphs, tool state, multimodal projectors, and sometimes a speculative-decoding model.

The budget looks like this:

flowchart TD
    T["Total available memory"] --> W["Model weights — mostly fixed"]
    T --> K["KV cache — grows with tokens and parallel slots"]
    T --> R["Runtime space — temporary working memory"]
    T --> H["Remaining headroom"]
    K --> C["Longer context uses more memory"]
    H --> F["Too little: spill, slow down, or crash"]

The rough budget is:

usable model memory = total memory - system use - weights - KV cache - runtime buffers

For a standard attention layer, KV-cache memory grows roughly like this:

tokens × layers × KV heads × head size × 2 (K + V) × bytes per value × parallel slots

The exact number depends heavily on the architecture. Grouped-query attention reduces the number of KV heads. Sliding-window, recurrent, and hybrid attention can reduce how much history must stay resident. Quantized K and V caches reduce it again. Multiple concurrent slots multiply it.

DeepSeek-V4.1-Flash shows how much the architecture can change the equation. DeepSeek reports a global KV-cache footprint of 890 bytes per token—about one quarter of its previous Flash model—using shared sparse-attention state and FP4 KV caching. A million-token context is still expensive, but it no longer has to scale like a conventional dense Transformer.

The important part is simpler: context is not free, and its cost grows with the context window you allocate.

This matters in llama.cpp because the configured context size reserves cache capacity. If I start a server with -c 262144, I am budgeting for a 262K-token cache even when the first prompt is tiny. On a 16 GB card, dropping to -c 32768 can be the difference between a useful 27B model and an out-of-memory error.

My starting rule is to leave at least:

  • 1.5–3 GB free on 4–12 GB devices
  • 4–8 GB free on 16–32 GB devices
  • 12–24 GB free on 48–128 GB devices
  • 10–20% free on larger unified-memory or multi-GPU systems

Those are planning numbers, not laws. The real test is peak memory after loading the model, ingesting a representative repository context, and completing several agent steps.

What changes at each tier

4–12 GB: a sharp assistant

Four-billion-parameter models are useful when the task is contained. Spark-X2.5-4B is small, agent-aware, and supports a huge theoretical context through a hybrid attention design. NeoHorse starts from Qwen3.5-4B and is tuned for tool use and coding. Neither turns 4 GB into a repo-scale coding agent.

At 8 GB, Ornith-1.5-9B is the first model in this list that starts to feel like a junior pair programmer instead of autocomplete. The limitation is not only intelligence. A roughly 5.75 GB model leaves only a little over 2 GB for context and runtime work.

The smart 12 GB move is precision, not parameter count. Keep the 9B model and move from 4-bit toward 8-bit. You get steadier output without creating a system that falls over when the conversation becomes interesting.

16–32 GB: the developer sweet spot

This is the range I find most interesting because smart quantization changes what the hardware can do.

The Qwen3.8-27B GSQ-RCO build assigns more bits to sensitive tensors and fewer bits to the ones that tolerate compression. Its recommended IQ3_S file is 11.8 GB, and its model card reports matching the BF16 model on AIME25 and LiveCodeBench v6. That makes a 27B model realistic on a 16 GB RTX 5080—as long as I do not spend the remaining 4 GB pretending context has no cost.

At 24 GB, a stronger 4-bit or low-5-bit build leaves enough room for a long file, a real conversation, and multi-file changes. At 32 GB, the choice is between a high-quality 27B quant and a sparse model such as Ornith-1.5-35B-A3B. I would default to the higher-precision Qwen. The Ornith model activates only about 3B parameters per token, which makes it fast, but sparse compute does not make its weights disappear.

48–128 GB: an agent, not a chat box

Qwen3-Coder-Next is the workhorse here. It has 80B total parameters, activates about 3B per token, and supports a native 262K context. More important, it was built for the loop that matters: inspect, edit, test, read the failure, and try again.

A compact third-party quant can fit around 38 GB on a 48 GB device. The official Q4_K_M is 48.4 GB, so I would put that on a 64 GB system, not a 48 GB one. Q6_K is 65.5 GB and Q8_0 is 84.8 GB, which gives the 80 GB and 96 GB tiers clean upgrade paths.

At 128 GB, Qwen3.8-Flash-Next becomes interesting. Its architecture includes components that do not behave like ordinary dense weights, and its 4-bit-class builds land around 94–111 GB. That is a better daily-driver fit than forcing GLM-5.3-Flash into the same 128 GB at an aggressive 3-bit quant with almost no margin.

192–512 GB: architecture beats the headline number

The large-memory tier is where parameter count becomes actively misleading.

GLM-5.3-Flash is a 320B mixture-of-experts model with 18B active parameters. MiniMax-M3 is roughly 428B total with 23B active. Both weights still have to live somewhere, but only a fraction participate in each token. That can make the larger sparse model faster and more practical than a smaller dense model on the same machine.

At 256 GB, MiniMax-M3's roughly 208 GB IQ4_XS build is the configuration I would want to use all day. It leaves room for context and runtime work while retaining the knowledge capacity of a much larger model.

At 512 GB, GLM-5.3 Q4 at roughly 467 GB is the trophy configuration. It fits. It also leaves less than 10% of the machine free before the context becomes serious. The roughly 343 GB 3-bit build is the configuration I would actually start with.

The most interesting experiment in the video is DeepSeek-V4.1-Flash through DwarfStar. The model includes 196B parameters of sparsely accessed conditional memory. DwarfStar keeps that colder lookup structure on SSD and streams pieces when needed, while the hot weights stay in fast memory. That can make a compressed 552B-backbone model possible on a 128 GB Mac, but it is still an experimental path, not something I would put under a deadline.

VRAM and unified memory are not interchangeable

A 128 GB unified-memory machine and a system with 128 GB of dedicated VRAM can both hold the same weights. They will not run them the same way.

Unified memory is shared with the operating system and every other process. Dedicated VRAM sits next to the accelerator and usually has much higher bandwidth. Multi-GPU systems add another variable: the links between cards. Once a model is split across devices, interconnect speed can matter as much as the memory on the spec sheet.

The same warning applies to CPU offload. Moving layers into system RAM may make the model fit, but every trip across PCIe can leave an expensive GPU waiting for data. Capacity solves the loading problem. Bandwidth determines how the system feels.

My actual rule

I do not choose the largest model that fits. I choose the smallest model that can reliably complete my real task, then spend the remaining memory on context and precision.

For most developers, that means a good 9B model on 8–12 GB, Qwen3.8-27B on 16–32 GB, or Qwen3-Coder-Next once 64 GB is available. Everything above that is real and useful, but the price rises much faster than the difference in day-to-day coding.

Fit is a specification. Helpfulness is a test.

Sources