AI Model Files Explained
A base model, an instruct fine-tune, a LoRA adapter, a Safetensors checkpoint, and a GGUF file describe different layers of the AI model stack. Here is how they fit together and which one to download.
Downloading a local model should be simple. Pick a model, download the file, and run it.
Then I open a model page and find a base model, an instruct version, three community fine-tunes, an adapter_model.safetensors, a dozen GGUF quantizations, and a warning that the wrong chat template will break everything.
The problem is not that AI has too many model types. The problem is that the word model is used for several different things:
- The architecture—the design of the neural network
- The checkpoint—the learned weights placed into that architecture
- The training stage—base, instruct, aligned, specialized, or distilled
- The artifact—full weights, a LoRA adapter, or a merged model
- The file format—Safetensors, GGUF, ONNX, MLX, or something else
- The precision—BF16, FP8, Q4, and so on
Those are separate axes. A model can be an instruct-tuned, decoder-only, 27B checkpoint stored as BF16 Safetensors. The same checkpoint can be converted into a Q4 GGUF without becoming a different architecture or learning anything new.
Once I separate those layers, the model page starts making sense.
The short version
| Term | What it describes | What it does not tell me |
|---|---|---|
| Architecture | The network's structure and operations | Which learned weights are loaded |
| Checkpoint | A saved set of learned weights | Whether it is packaged for my runtime |
| Pretrained / base | The broad starting checkpoint | Whether it follows chat instructions well |
| Fine-tune | A checkpoint trained further for a behavior, domain, or task | Whether all weights or only adapters were trained |
| Instruct / chat | A fine-tune designed to follow requests and dialogue | Whether it is best for completion or embedding tasks |
| LoRA | A small set of learned weight updates for a specific base | A complete standalone model |
| Merged model | Base weights with an adapter applied permanently | The original adapter's flexibility |
| Safetensors | A safe, fast tensor-storage format | The architecture, purpose, or quantization by itself |
| GGUF | A self-describing format built for efficient local inference | That the model must be quantized |
| Quantization | A lower-precision representation of weights | The model's training stage or task |
The shortest useful rule is:
Pretraining and fine-tuning change what the model knows or how it behaves. LoRA changes how the fine-tune is stored. Safetensors and GGUF change how the weights are packaged. Quantization changes how precisely they are represented.
A model is not a file
An architecture is the blueprint: the number and type of layers, attention design, hidden dimensions, expert routing, positional encoding, and every operation needed to transform input tokens into output scores.
A checkpoint is the set of learned numbers loaded into that blueprint.
The distinction is similar to code and data. The architecture defines what calculations exist. The checkpoint supplies the values those calculations use.
This is why renaming a file never converts a model. A runtime has to understand the architecture, map every tensor to the correct operation, load the tokenizer, preserve special tokens, and reproduce the expected chat formatting. The weights are essential, but they are not the entire application.
A complete model package may include:
- Model weights
- An architecture configuration
- A tokenizer and vocabulary
- Special-token definitions
- A generation configuration
- A chat template
- Quantization metadata
- Multimodal projectors or encoders
- Custom model code
The extension tells me how some of that data is stored. It does not tell me what the model is good at.
The model lifecycle
Most downloadable models are descendants of an expensive pretrained checkpoint. They branch into specialized versions, adapters, and inference formats after that.
flowchart TD
D["Large training corpus"] --> P["Pretraining"]
P --> B["Base checkpoint"]
B --> C["Continued pretraining"]
B --> S["Instruction fine-tuning"]
S --> A["Preference alignment"]
B --> L["LoRA or QLoRA"]
C --> E["Full or adapter checkpoint"]
A --> E
L --> E
E --> M["Optional merge"]
M --> F["Safetensors or GGUF export"]
F --> Q["Optional quantization"]
This diagram is simplified. A real training recipe can repeat stages, combine multiple datasets, distill from another model, or apply preference training to an adapter. The important part is the direction: format conversion happens after the model has learned its behavior. Converting Safetensors to GGUF is not training.
Pretrained and base models
Pretraining is the large, expensive stage that gives a model its general capabilities. For a causal language model, the training objective is usually some version of predicting the next token from the tokens before it. The labels come from the text itself, which is why this is called self-supervised learning.
The resulting checkpoint is usually called a pretrained model, foundation model, or base model.
A base model has learned language, facts, patterns, code, and relationships from its training corpus. It has not necessarily learned to behave like a helpful chat assistant. If I ask a base model a question, it may continue the question, imitate a document, or produce a plausible next passage instead of answering directly.
Base models are useful for:
- Further fine-tuning
- Research and evaluation
- Text completion
- Domain adaptation
- Building a custom assistant from a clean starting point
For ordinary chat or agentic coding, I normally choose an instruct model rather than the base unless I intend to train it myself.
Continued pretraining is not the same as instruction tuning
Continued pretraining, sometimes called domain-adaptive pretraining, takes a pretrained model and continues the original language-modeling objective on more data.
If I continue pretraining on source code, legal text, medical literature, or another language, I am trying to change the model's underlying distribution and knowledge. The data does not need to be written as question-and-answer pairs.
Continued pretraining is useful when the base model has not seen enough of the target domain. It can also cause catastrophic forgetting if the new data is too narrow or the training is too aggressive: the model becomes better at the new distribution while losing some of its broader ability.
Instruction tuning teaches a different lesson. It trains the model on prompts and desired responses so it learns how to follow requests. A model can know a great deal about Python after pretraining and still need instruction tuning to reliably respond with only a patch, call a tool in the required JSON schema, or stop after completing the task.
Knowledge and behavior overlap, but they are not the same target.
Fine-tuning is an umbrella term
Fine-tuning means continuing training from an existing checkpoint. It does not identify one specific technique.
The major categories are:
| Fine-tuning stage | Primary goal | Typical data |
|---|---|---|
| Continued pretraining | Add domain or language exposure | Raw text or code |
| Supervised fine-tuning (SFT) | Teach desired input-output behavior | Prompt and response examples |
| Instruction tuning | Make the model follow diverse requests | Instruction-response conversations |
| Preference tuning | Prefer one answer or behavior over another | Chosen/rejected response pairs or reward signals |
| Task fine-tuning | Optimize for a narrow task | Labeled examples |
| Distillation | Transfer behavior from a teacher model | Teacher outputs, logits, or generated traces |
Supervised fine-tuning, or SFT, trains against known target outputs. The model is penalized when its predicted tokens differ from the desired response.
Preference alignment goes beyond imitation. Methods such as RLHF, RLAIF, and DPO use human or AI preferences to make some responses more likely than others. This can shape helpfulness, safety, style, verbosity, refusal behavior, and reasoning strategy.
A model named Instruct, Chat, or Assistant has usually gone through SFT and possibly one or more preference stages. The exact recipe should be in the model card; the name alone is not a technical specification.
Full fine-tuning versus parameter-efficient fine-tuning
In a full fine-tune, the training process can update all model weights. This provides maximum freedom but requires substantial memory for gradients, optimizer states, activations, and often higher-precision master weights.
Parameter-efficient fine-tuning, or PEFT, freezes most of the base model and trains a much smaller set of parameters. LoRA is the most common example, but it is not the only PEFT method.
| Approach | What is trained | Storage for each variant | Main tradeoff |
|---|---|---|---|
| Full fine-tune | Most or all base weights | Another full checkpoint | Maximum flexibility, highest training cost |
| LoRA | Small low-rank update matrices | A compact adapter | Efficient and portable, dependent on the base |
| Prompt / prefix tuning | Learned virtual prompt representations | Very small adapter | Cheap, but less flexible for some tasks |
| Partial or layer tuning | Selected layers or parameters | Between adapter and full checkpoint | More capacity with more cost |
PEFT reduces trainable parameters. It does not make the frozen base disappear. I still need the base model in memory during training and inference.
What a LoRA actually is
LoRA stands for Low-Rank Adaptation. Instead of updating a large weight matrix directly, LoRA learns two much smaller matrices whose product represents an update to that weight.
In simplified form:
effective weight = base weight + scaled LoRA update
W' = W + scale × B × A
The base matrix W stays frozen. Only A and B are trained.
The word rank describes the inner dimension of those smaller matrices. A higher rank gives the adapter more capacity but increases training cost and adapter size. lora_alpha controls the update's scaling, and target_modules determines which parts of the network receive LoRA layers.
This is why two LoRAs for the same model can differ dramatically. One may target only attention projections at rank 8. Another may target every linear layer at rank 64. Both are “LoRAs,” but they do not have the same capacity, size, or effect.
LoRA works well because many useful model adaptations can be represented with a much lower-rank update than the full weight matrix. The original LoRA paper freezes the pretrained weights and injects trainable rank-decomposition matrices, reducing the number of parameters that need to be trained.
A LoRA is normally not a complete model
A LoRA adapter might contain:
adapter_model.safetensors
adapter_config.json
The Safetensors file contains the learned adapter tensors. The configuration records information such as the PEFT method, rank, scaling, target modules, and often the expected base model.
To use it, I need:
- The correct base model architecture
- The correct base checkpoint or a compatible revision
- A compatible tokenizer and configuration
- A runtime that knows how to apply that adapter
An adapter trained for one 27B model cannot be assumed to work on another 27B model. Matching parameter count is not enough. Tensor names, shapes, layer counts, architecture revisions, and tokenizer choices must agree.
This matters just as much for image-generation LoRAs. A LoRA trained for one diffusion-model family is not automatically compatible with another merely because both generate images.
LoRA at runtime versus a merged model
There are two common ways to deploy a LoRA.
Load the base and adapter separately
The runtime loads the base model, loads the adapter, and applies the LoRA update during inference.
This is useful when I want to:
- Switch among several adapters
- Combine adapters supported by the runtime
- Keep one shared base model on disk
- Preserve the original adapter for more training
The cost is added runtime complexity and, depending on the engine and targeted layers, some inference overhead.
Merge the adapter into the base
Merging calculates the effective weights and saves a new standalone checkpoint. The resulting model no longer needs the LoRA files at inference time.
This is useful when I want to:
- Convert the result into GGUF
- Quantize one final deployment model
- Use a runtime without PEFT support
- Eliminate adapter application overhead
Merging gives up easy adapter switching and usually creates another full-size checkpoint. It can also be lossy or unsupported when performed directly against aggressively quantized weights. My safest deployment path is to merge into a high-precision base, verify the result, and then quantize the merged model.
Some runtimes can apply a LoRA directly to a quantized GGUF base. That is convenient for experimentation, but it is a different workflow from creating one permanently merged and quantized file.
QLoRA is a training technique, not a special model family
QLoRA combines a frozen quantized base model with trainable LoRA adapters. The original QLoRA work loads the base in 4-bit NormalFloat, performs computation at a higher precision, and backpropagates through the frozen quantized model into the LoRA weights.
The key idea is:
4-bit frozen base + trainable LoRA adapters = much lower fine-tuning memory
QLoRA does not mean the adapter itself is a Q4 model. The adapter weights are normally stored separately in a higher-precision tensor format. It also does not mean the final merged model must remain 4-bit. I can merge the learned update into a higher-precision base checkpoint and then choose a deployment precision.
QLoRA makes fine-tuning far more accessible, but training still needs memory for activations, gradients for the adapter, optimizer state, batches, and sequence length. Inference memory and fine-tuning memory are not interchangeable.
Safetensors: a safe box for tensors
Safetensors is a storage format for tensors—the multidimensional arrays that hold model weights and other learned values. Hugging Face describes it as a safe alternative to pickle with fast, zero-copy loading.
The security distinction matters. Python pickle can reconstruct arbitrary Python objects and may execute code during deserialization. Safetensors stores a restricted header plus raw tensor data; the file is not a general Python-object program.
That does not make every repository containing Safetensors automatically safe. A model repo may include custom Python code, and loading it with options such as trust_remote_code=True can execute that code. Safetensors makes the weight file safer; it does not bless every accompanying file or runtime action.
Safetensors is especially useful in training and GPU inference ecosystems because it:
- Preserves named tensors and their data types
- Supports memory mapping and efficient slices
- Avoids pickle's arbitrary-code deserialization
- Works well with sharded checkpoints
- Integrates with Transformers, Diffusers, PEFT, MLX, and other tools
One Safetensors file may not be a complete model
Large checkpoints are often split into files such as:
model-00001-of-00008.safetensors
model-00002-of-00008.safetensors
...
model.safetensors.index.json
The index maps tensor names to shards. The repository also needs configuration and tokenizer files. Copying only one shard will not produce a smaller working model; it produces an incomplete checkpoint.
A typical Transformers-style repository may look like:
config.json
generation_config.json
tokenizer.json
tokenizer_config.json
special_tokens_map.json
chat_template.jinja
model-00001-of-00008.safetensors
...
model.safetensors.index.json
The weights are in Safetensors. The model definition and behavior depend on the rest of the package.
GGUF: an inference package, not a synonym for Q4
GGUF is a binary format from the GGML ecosystem designed to package tensors and model metadata for efficient inference. It is most closely associated with llama.cpp and the many applications built around it.
A GGUF can include:
- Architecture metadata
- Model tensors and their data types
- Tokenizer vocabulary and token metadata
- Context and positional-encoding parameters
- Quantization information
- A chat template and other metadata
The format is designed to be self-describing enough for a compatible runtime to map and execute the model efficiently. It supports memory mapping, which lets the operating system map file regions instead of copying the entire file through a traditional read path.
GGUF is popular for local inference because llama.cpp can run across CPUs and multiple accelerator backends, split layers between CPU and GPU, and support a wide range of block quantizations.
But GGUF does not mean 4-bit. A GGUF can contain F32, F16, BF16, Q8, Q6, Q5, Q4, Q3, Q2, or a mixture of tensor types. The extension tells me the container. The quant name and file size tell me the weight representation.
GGUF also does not guarantee universal compatibility. A runtime must implement the specific model architecture, tensor types, tokenizer behavior, and metadata version. A brand-new model may have valid GGUF files before my installed runtime knows how to execute them.
Safetensors versus GGUF
They overlap, but they are optimized for different parts of the workflow.
| Question | Safetensors | GGUF |
|---|---|---|
| Primary role | Safe tensor storage for training and framework-native checkpoints | Self-describing package for efficient local inference |
| Common ecosystem | Transformers, PyTorch, Diffusers, PEFT, MLX | llama.cpp, LM Studio, KoboldCpp, compatible local apps |
| Typical contents | Named tensors; companion files provide config and tokenizer | Tensors plus substantial model and tokenizer metadata |
| Quantization | Can store original or quantized tensors supported by the framework | Commonly used with GGML block quantizations; can also store higher precision |
| Fine-tuning | Natural source format for training and adapters | Primarily an inference format, though adapters can also be converted |
| Sharding | Common for large models | Often one file, but large GGUF models can be split |
| Hardware path | Depends on PyTorch/framework kernels | Designed for portable CPU/GPU inference backends |
| Best choice | Training, fine-tuning, framework-native GPU serving | Simple local deployment and flexible hardware offload |
Neither format is inherently smarter or higher quality. If both contain equivalent weights at equivalent precision and both runtimes implement the model correctly, the model has the same learned capability. Differences come from quantization, kernels, prompt formatting, sampler settings, and implementation details—not the filename's personality.
Converting Safetensors to GGUF
The normal llama.cpp deployment path is:
Transformers checkpoint
↓
BF16/F16 GGUF conversion
↓
GGUF quantization
↓
llama.cpp-compatible runtime
The converter reads the architecture configuration, maps tensor names and layouts, imports tokenizer data, writes GGUF metadata, and serializes the weights. A quantizer can then create smaller variants from the high-precision GGUF.
This is why I prefer to keep the original high-precision checkpoint or at least one high-quality intermediate. Re-quantizing an already low-bit model compounds approximation error. A Q4 file is a deployment artifact, not the ideal master copy for producing every future format.
Conversion can fail or silently behave badly when:
- The architecture is not supported yet
- Tensor layouts changed in a new model revision
- The tokenizer or special tokens are missing
- A multimodal projector was not converted
- The chat template is wrong or absent
- The model uses an unsupported quantization or custom layer
“Converted successfully” only means the file was written. I still test generation, special tokens, tool calling, long context, and task quality.
Other model formats I run into
Safetensors and GGUF cover much of my local workflow, but they are not the only formats.
| Format | Where it fits |
|---|---|
.bin, .pt, .pth | PyTorch checkpoints; may use pickle and require more trust when loading |
| ONNX | Portable computation graph for inference across ONNX runtimes |
| EXL2 | GPU-focused quantization format used by ExLlamaV2-style runtimes |
| GPTQ / AWQ checkpoints | Quantized weight layouts used by supported GPU inference engines |
| MLX | Apple-silicon-focused model packages for the MLX ecosystem |
| TensorRT engines | Highly optimized, hardware- and configuration-specific NVIDIA inference artifacts |
| Adapter Safetensors | Small PEFT or LoRA updates that require a compatible base |
Some names in that table describe a file format, some a quantization method, and some an execution ecosystem. The industry does not maintain a clean taxonomy, so I always ask two questions: What is stored? Which runtime is expected to load it?
Model types by architecture and task
The training and file-format labels sit on top of another distinction: what kind of network and output the model was designed to produce.
Decoder-only causal language models
These predict the next token from previous tokens. Most modern chat, coding, and agentic LLMs belong here. They are naturally suited to open-ended generation.
Encoder-only models
These build representations from the full input rather than generating long continuations one token at a time. They are common for classification, named-entity recognition, embeddings, and some rerankers.
Encoder-decoder models
The encoder reads an input sequence and the decoder generates a target sequence. Translation and summarization are classic uses.
Embedding models
An embedding model turns text, images, or other inputs into vectors. Similar items should land near one another in vector space. These models power semantic search and retrieval; they are not chat models even when both use Transformer components.
Rerankers
A reranker scores a query-document pair more precisely than a basic embedding similarity search. A common retrieval pipeline uses embeddings to find candidates quickly and a reranker to reorder the best few.
Vision-language and multimodal models
A VLM combines a language model with a vision encoder, projector, or integrated multimodal architecture. The text weights alone may load while image understanding remains unavailable because a required projector or vision tower is missing.
Diffusion and other generative-media models
Image, audio, and video generators may use Transformers, diffusion components, VAEs, text encoders, and LoRAs. They can use Safetensors or GGUF too, but their pipeline and runtime requirements differ from an autoregressive LLM.
The extension does not define the task. A .safetensors file could be an LLM shard, a LoRA, a diffusion transformer, a VAE, a text encoder, or an embedding model.
Base, instruct, reasoning, coding, and agentic are training labels
Model names often include behavior-oriented labels:
| Label | What I expect | What I still verify |
|---|---|---|
| Base | General pretrained completion model | Training data, license, context |
| Instruct / Chat | Request following and conversation | Chat template and alignment quality |
| Reasoning | More deliberate multi-step behavior | Token cost, mode controls, benchmark relevance |
| Coder | More code and software-task training | Languages, fill-in-the-middle, tool use |
| Agentic | Tool use and multi-step environment interaction | Harness compatibility and recovery behavior |
| Embedding | Vector output for retrieval or similarity | Vector size, pooling, supported languages |
| Reranker | Query-document relevance scoring | Input length and serving cost |
| Abliterated / uncensored | Alignment or refusal behavior was modified | Capability loss, provenance, and safety implications |
These labels are useful search terms, not guarantees. A coder model can be poor at editing a real repository. An agentic model can produce perfect tool-call syntax but fail to recover from a test error. I treat the model card and my own evaluations as the specification.
How to read a model repository name
Suppose I see:
Example-27B-Instruct-LoRA
Example-27B-Instruct-AWQ
Example-27B-Instruct-GGUF
Example-27B-Instruct-Q4_K_M.gguf
I read it from left to right:
Exampleidentifies the model family or architecture.27Bgives the approximate parameter count.Instructdescribes the training stage or behavior.LoRA,AWQ, orGGUFdescribes the artifact, quantization, or ecosystem.Q4_K_Midentifies the specific quantized tensor recipe.
The LoRA repository probably does not contain full weights. The AWQ repository likely contains a complete quantized checkpoint for an AWQ-compatible loader. The GGUF repository may contain several downloadable quant levels. The final filename identifies one of those files.
Community naming is inconsistent, so I verify the contents instead of trusting the suffix.
The files I check before downloading
The largest file gets the attention, but these are just as important:
Model card
The model card should identify the base model, training recipe, intended uses, limitations, evaluation results, prompt format, and license. If a fine-tune does not clearly name its base, I consider that a deployment risk.
License
Open weights do not automatically mean unrestricted use. The base model, fine-tune dataset, adapter, and redistributed merge may have different conditions. Merging weights does not erase their licenses.
Configuration
config.json describes the architecture. An unsupported model_type or custom layer may require a newer runtime or repository code.
Tokenizer
Tokenizer files define how text becomes token IDs. A mismatch can corrupt output even if every model tensor loads correctly.
Chat template
The template formats system, user, assistant, and tool messages into the token pattern used during instruction training. The wrong template can cause broken role handling, repeated text, poor tool calls, or immediate end-of-sequence tokens.
Checksums and provenance
I prefer files traceable to the model author or a reputable converter. Exact hashes matter when a model is updated without a clear filename change.
My preferred workflows
I only want to run the model locally
I choose an instruct or specialized checkpoint, find a GGUF quant that fits with context headroom, confirm my runtime supports the architecture, and use the model's expected chat template.
I want maximum GPU serving performance
I start with the original Safetensors checkpoint and choose the quantization and engine supported by my GPU stack—BF16, FP8, AWQ, GPTQ, or another optimized path. GGUF is wonderfully portable, but it is not automatically the fastest format for every NVIDIA deployment.
I want to create a custom personality or domain behavior
I keep the original base or instruct checkpoint in Safetensors, train a LoRA or QLoRA adapter, evaluate the adapter against the exact base, and preserve both before merging.
I want one deployable GGUF from my LoRA
My clean path is:
exact base checkpoint
+
LoRA adapter
↓
merge at high precision
↓
export or convert to high-precision GGUF
↓
quantize the merged GGUF
↓
test prompt format and real tasks
That workflow gives me a reusable adapter, a high-quality merged master, and one or more deployment quants. It costs more disk space, but it avoids turning the smallest artifact into the only copy I can build from.
My actual rule
I do not ask, “Is GGUF better than Safetensors?” They solve different problems.
I ask:
- What architecture and task is this model built for?
- Is this a base, instruct, aligned, or specialized checkpoint?
- Is the download a complete model or only an adapter?
- What precision are the weights stored in?
- Which runtime is expected to load them?
- Are the tokenizer, template, license, and base-model revision clear?
Safetensors is where I prefer to keep trainable, framework-native source checkpoints and adapters. GGUF is where I prefer to package a practical local-inference build. LoRA is how I avoid duplicating an entire base model for every experiment. Merging is how I turn one successful experiment into a standalone deployment artifact.
The file is not the model. It is one layer of the model's history.
Sources
- Hugging Face LLM Course: How Transformers work
- Hugging Face Transformers documentation
- Hugging Face Safetensors documentation
- Hugging Face PEFT: LoRA
- Hugging Face PEFT checkpoint format
- GGUF specification
llama.cppHugging Face-to-GGUF converter- LoRA: Low-Rank Adaptation of Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Direct Preference Optimization