Dense vs. MoE Models Explained
A 30B model with 3B active parameters still needs room for its full weights. Here is how dense and mixture-of-experts models differ and what total parameters, routing, memory, and compute mean for local AI.
The first time I see a model labeled 30B-A3B, the obvious question is: am I running a 30-billion-parameter model or a 3-billion-parameter model?
Both numbers describe something real. Neither tells the whole story.
The first describes how many parameters the model contains. The second describes how many participate in processing one token. That distinction is the reason a mixture-of-experts model can have relatively low compute requirements and still need a substantial amount of memory.
It also explains why comparing model names by the biggest number can be misleading.
The short version
| Question | Dense model | Sparse mixture-of-experts model |
|---|---|---|
| How does a token move through the model? | Through the same attention and feed-forward blocks | Through shared components and a selected subset of expert blocks |
| What changes between tokens? | The activations; the network path is broadly fixed | The activations and the selected experts |
| Which parameter count matters for weight storage? | Total parameters | Total parameters, including inactive experts |
| Which count helps estimate per-token weight computation? | Broadly the total count | Active parameters, with architecture-specific qualifications |
| What is the main tradeoff? | Simpler execution, but more computation as the model grows | More parameter capacity per unit of compute, with routing and memory costs |
Dense models use a broadly fixed network path. Sparse MoE models choose which expert blocks run for each token.
The word sparse matters. This article is about the sparse MoE architectures commonly used in LLMs, where the runtime skips unselected experts. An MoE that evaluates every expert would not offer the same compute savings. Hugging Face's MoE explainer covers that distinction.
What a dense model actually does
A Transformer block has two major jobs: attention mixes information across tokens, and a feed-forward network transforms each token's representation.
In a dense model, tokens pass through the same feed-forward network at a given layer. Different inputs produce different activations, but there is no router choosing between alternative expert networks.
A simplified block looks like this:
Token representation
↓
Attention
↓
One feed-forward network
↓
Next layer
Residual connections and normalization are omitted here to keep the picture readable.
People often summarize this as “every parameter is active for every token.” That is useful shorthand, but it is not literal accounting: an embedding lookup, for example, selects a row rather than multiplying through the entire embedding table.
The useful distinction is the absence of conditional expert selection.
What changes in an MoE model
In many MoE Transformers, some or all feed-forward blocks are replaced with several alternative networks called experts. A small learned network called a router, or gate, chooses which experts process the token's current representation.
For a layer with eight experts and top-2 routing:
Token representation
↓
Attention
↓
Router
├── Expert 2 ──┐
└── Expert 6 ──┤
↓
Weighted combination
↓
Next layer
The other six experts do no feed-forward computation for that token at that layer. Their weights still exist.
Top-k means the router selects k experts. Top-2 selects two; top-8 selects eight. The choice happens at each MoE layer, so the next layer can select a different pair. The next token can take a different path too.
This is the mechanism used by Mixtral: eight feed-forward experts per layer, with two selected for each token.
An expert is not a separate chatbot
The name makes it easy to imagine a coding model, a math model, and a writing model sitting behind a dispatcher.
That is a different system.
In a typical Transformer MoE, each expert is a neural-network block inside the model. It receives a numerical representation and returns another numerical representation. It does not read the whole conversation and write its own answer.
The router operates on those representations, rather than a human-readable instruction such as “send this prompt to the Python specialist.” Expert behavior emerges from training, and the resulting specialization need not match subjects I would name.
Mixtral's authors found routing patterns associated with syntax rather than a clean separation by subject. That is a useful correction to the mental picture of a room full of specialist chatbots. Mixtral routing analysis
Total parameters versus active parameters
Total parameters count the stored model weights, including all experts.
Active parameters describe the weights involved in the selected computation path for one token. Counting conventions can differ, so the model card matters more than the shorthand in the name.
Here are three concrete examples from published architectures:
| Model | Architecture | Total parameters | Active parameters per token | Experts selected per MoE layer |
|---|---|---|---|---|
| Qwen3-32B | Dense | 32.8B | Broadly the dense network | No expert router |
| Qwen3-30B-A3B | MoE | 30.5B | 3.3B | 8 of 128 |
| Mixtral 8x7B | MoE | About 47B | About 13B | 2 of 8 |
The Qwen figures come from the Qwen3-32B model card and the Qwen3-30B-A3B model card. The Mixtral figures come from its release announcement.
These are architecture examples, not a ranking of the best models to download.
Notice that 8x7B does not mean eight complete 7B chatbots or exactly 56B stored parameters. Components outside the expert feed-forward blocks are shared, so multiplying the numbers in the name gives the wrong total.
Likewise, A3B is a rounded naming convention. For this Qwen checkpoint, the documented active count is 3.3B.
Why selecting 8 of 128 experts is not 1/16 of everything
Only the routed portion receives that reduction. Attention, the router, and other shared components still run.
A simplified accounting model is:
total parameters ≈ S + E × P
active parameters ≈ S + k × P
S = shared parameters
E = number of available experts
P = parameters in each expert
k = experts selected per token
For multiple layers, add their contributions. This illustration assumes equally sized experts and treats shared parameters as active for counting purposes; it is not an exact FLOP formula.
Some architectures also have shared experts that always run alongside the selected experts. DeepSeek-V3 uses this arrangement, together with fine-grained routed experts and a load-balancing strategy. Shared experts belong in the always-active part of the estimate. DeepSeek-V3 technical report
The fraction k / E describes expert selection. It does not describe the fraction of the entire model's work.
Memory follows the total count
This is the part I check before downloading a model.
For an ordinary deployment, the full set of weights must be available somewhere: accelerator memory, unified memory, system RAM, or a combination. The router can choose different experts on the next token, so I cannot permanently discard the experts that were inactive on the previous one.
The basic storage calculation is the same as in LLM Quantization Explained:
raw weight storage = total parameters × bits per weight ÷ 8
Using Qwen3-30B-A3B's documented 30.5B total parameters:
| Representation | Raw storage for all 30.5B parameters | Incorrect estimate using only 3.3B active |
|---|---|---|
| BF16, 16 bits | 61 GB | 6.6 GB |
| Hypothetical exact 4-bit weights | 15.25 GB | 1.65 GB |
These are decimal gigabytes calculated from parameter counts. They are storage estimates, not download sizes or measured runtime requirements. Real quants include scales, metadata, and tensors stored at different precisions.
I also need room for the KV cache, activations, runtime buffers, and the rest of the machine.
An MoE with 3.3B active parameters therefore does not have the memory footprint of a dense 3.3B model. Expert offloading or caching can reduce accelerator residency, but moving or computing those weights elsewhere adds costs that need to be measured.
Compute savings do not guarantee a speed multiplier
MoE's main appeal is that it can increase learned parameter capacity without evaluating every expert for every token. Conditional computation was a central motivation behind Switch Transformers.
That gives active parameter count a useful role when estimating weight computation. It does not turn 30.5 / 3.3 into a guaranteed 9.2× speedup.
The runtime still has to select experts, dispatch tokens, execute the selected networks efficiently, and combine their outputs. Shared attention still costs time. Small expert workloads may use hardware less efficiently than one large dense operation.
When experts are spread across GPUs, dispatch also involves communication. Expert parallelism assigns different experts to different devices; the tokens must travel to the devices holding their selected experts, and the results must come back. Hardware interconnects and load balance become part of the performance story. vLLM's expert-parallel deployment guide describes these serving considerations.
Prefill and decode still need separate measurements
The distinction from Prefill vs. Decode applies here too.
During prefill, many prompt tokens are processed together. Each token selects only a few experts, but the prompt as a whole may touch many or all of them. How the runtime groups those tokens into expert operations affects utilization.
During decode for one conversation, relatively small amounts of work reach each selected expert. Weight bandwidth, kernel overhead, attention, and any offloading can become substantial costs.
Serving many conversations changes the picture again: different tokens can select different experts, and the backend can group their work. Higher aggregate throughput does not automatically mean lower latency for my one request.
I want measured prompt-processing speed, time to first token, and output speed at the context lengths I use. An active parameter count cannot replace those measurements.
Training has its own tradeoffs
The router and experts are learned parts of the model. If a router sends most tokens to a few experts, those experts receive most of the work while others are underused. That creates both a learning problem and an execution bottleneck.
MoE training therefore needs mechanisms to distribute work. Switch Transformers uses a balancing objective; DeepSeek-V3 describes a different approach that adjusts routing with expert-specific biases, alongside a sequence-level balancing loss.
The specific method varies. The shared problem is making the extra parameter capacity useful without letting routing become a bottleneck. Switch Transformers, DeepSeek-V3
For fine-tuning, I still need to understand which parameters an adapter targets and how the training data exercises the experts. “Supports LoRA” is only the start of the configuration.
Which one is better for local AI?
I would not choose from the architecture label alone.
A dense model can be attractive when it fits comfortably, has mature runtime support, and gives me reliable results. Its simpler execution path makes performance easier to reason about.
An MoE can be attractive when I have enough memory for the full checkpoint and the backend executes its experts efficiently. Its extra parameter capacity may deliver better results for a given amount of per-token compute.
Neither total parameters nor active parameters establishes an equivalent dense model's quality. Training data, post-training, quantization, and the task all matter. A 30B-A3B model is neither guaranteed to behave like a dense 30B nor limited to the capability of a dense 3B.
For my local coding workflow, I compare:
- The actual checkpoint size, with room left for context and runtime buffers.
- Support for the exact architecture and quantization, on the backend I intend to use.
- Prefill and decode at realistic prompt sizes, including repeated tool results and repository context.
- Results on real tasks, such as editing files, following instructions, calling tools, and recovering from errors.
- Time to a useful result, including retries and any additional reasoning tokens.
Fast tokens are helpful. A completed task is the outcome I care about.
My actual rule
When I read an MoE model name, I keep three measurements separate:
Total parameters → estimate full weight storage
Active parameters → understand selected per-token computation
Measured results → decide whether it is useful on my machine
MoE makes a larger pool of learned weights available while using only part of the expert pool at a time. That is a powerful architectural tradeoff.
It still leaves me with the practical work of fitting the model, running it efficiently, and checking whether its answers are good.