← 05 / Writing
2026-09-20//local-ai21 min readLIVE

AI Model Files Explained

A base model, an instruct fine-tune, a LoRA adapter, a Safetensors checkpoint, and a GGUF file describe different layers of the AI model stack. Here is how they fit together and which one to download.

Downloading a local model should be simple. Pick a model, download the file, and run it.

Then I open a model page and find a base model, an instruct version, three community fine-tunes, an adapter_model.safetensors, a dozen GGUF quantizations, and a warning that the wrong chat template will break everything.

The problem is not that AI has too many model types. The problem is that the word model is used for several different things:

  • The architecture—the design of the neural network
  • The checkpoint—the learned weights placed into that architecture
  • The training stage—base, instruct, aligned, specialized, or distilled
  • The artifact—full weights, a LoRA adapter, or a merged model
  • The file format—Safetensors, GGUF, ONNX, MLX, or something else
  • The precision—BF16, FP8, Q4, and so on

Those are separate axes. A model can be an instruct-tuned, decoder-only, 27B checkpoint stored as BF16 Safetensors. The same checkpoint can be converted into a Q4 GGUF without becoming a different architecture or learning anything new.

Once I separate those layers, the model page starts making sense.

The short version

TermWhat it describesWhat it does not tell me
ArchitectureThe network's structure and operationsWhich learned weights are loaded
CheckpointA saved set of learned weightsWhether it is packaged for my runtime
Pretrained / baseThe broad starting checkpointWhether it follows chat instructions well
Fine-tuneA checkpoint trained further for a behavior, domain, or taskWhether all weights or only adapters were trained
Instruct / chatA fine-tune designed to follow requests and dialogueWhether it is best for completion or embedding tasks
LoRAA small set of learned weight updates for a specific baseA complete standalone model
Merged modelBase weights with an adapter applied permanentlyThe original adapter's flexibility
SafetensorsA safe, fast tensor-storage formatThe architecture, purpose, or quantization by itself
GGUFA self-describing format built for efficient local inferenceThat the model must be quantized
QuantizationA lower-precision representation of weightsThe model's training stage or task

The shortest useful rule is:

Pretraining and fine-tuning change what the model knows or how it behaves. LoRA changes how the fine-tune is stored. Safetensors and GGUF change how the weights are packaged. Quantization changes how precisely they are represented.

A model is not a file

An architecture is the blueprint: the number and type of layers, attention design, hidden dimensions, expert routing, positional encoding, and every operation needed to transform input tokens into output scores.

A checkpoint is the set of learned numbers loaded into that blueprint.

The distinction is similar to code and data. The architecture defines what calculations exist. The checkpoint supplies the values those calculations use.

This is why renaming a file never converts a model. A runtime has to understand the architecture, map every tensor to the correct operation, load the tokenizer, preserve special tokens, and reproduce the expected chat formatting. The weights are essential, but they are not the entire application.

A complete model package may include:

  • Model weights
  • An architecture configuration
  • A tokenizer and vocabulary
  • Special-token definitions
  • A generation configuration
  • A chat template
  • Quantization metadata
  • Multimodal projectors or encoders
  • Custom model code

The extension tells me how some of that data is stored. It does not tell me what the model is good at.

The model lifecycle

Most downloadable models are descendants of an expensive pretrained checkpoint. They branch into specialized versions, adapters, and inference formats after that.

flowchart TD
    D["Large training corpus"] --> P["Pretraining"]
    P --> B["Base checkpoint"]
    B --> C["Continued pretraining"]
    B --> S["Instruction fine-tuning"]
    S --> A["Preference alignment"]
    B --> L["LoRA or QLoRA"]
    C --> E["Full or adapter checkpoint"]
    A --> E
    L --> E
    E --> M["Optional merge"]
    M --> F["Safetensors or GGUF export"]
    F --> Q["Optional quantization"]

This diagram is simplified. A real training recipe can repeat stages, combine multiple datasets, distill from another model, or apply preference training to an adapter. The important part is the direction: format conversion happens after the model has learned its behavior. Converting Safetensors to GGUF is not training.

Pretrained and base models

Pretraining is the large, expensive stage that gives a model its general capabilities. For a causal language model, the training objective is usually some version of predicting the next token from the tokens before it. The labels come from the text itself, which is why this is called self-supervised learning.

The resulting checkpoint is usually called a pretrained model, foundation model, or base model.

A base model has learned language, facts, patterns, code, and relationships from its training corpus. It has not necessarily learned to behave like a helpful chat assistant. If I ask a base model a question, it may continue the question, imitate a document, or produce a plausible next passage instead of answering directly.

Base models are useful for:

  • Further fine-tuning
  • Research and evaluation
  • Text completion
  • Domain adaptation
  • Building a custom assistant from a clean starting point

For ordinary chat or agentic coding, I normally choose an instruct model rather than the base unless I intend to train it myself.

Continued pretraining is not the same as instruction tuning

Continued pretraining, sometimes called domain-adaptive pretraining, takes a pretrained model and continues the original language-modeling objective on more data.

If I continue pretraining on source code, legal text, medical literature, or another language, I am trying to change the model's underlying distribution and knowledge. The data does not need to be written as question-and-answer pairs.

Continued pretraining is useful when the base model has not seen enough of the target domain. It can also cause catastrophic forgetting if the new data is too narrow or the training is too aggressive: the model becomes better at the new distribution while losing some of its broader ability.

Instruction tuning teaches a different lesson. It trains the model on prompts and desired responses so it learns how to follow requests. A model can know a great deal about Python after pretraining and still need instruction tuning to reliably respond with only a patch, call a tool in the required JSON schema, or stop after completing the task.

Knowledge and behavior overlap, but they are not the same target.

Fine-tuning is an umbrella term

Fine-tuning means continuing training from an existing checkpoint. It does not identify one specific technique.

The major categories are:

Fine-tuning stagePrimary goalTypical data
Continued pretrainingAdd domain or language exposureRaw text or code
Supervised fine-tuning (SFT)Teach desired input-output behaviorPrompt and response examples
Instruction tuningMake the model follow diverse requestsInstruction-response conversations
Preference tuningPrefer one answer or behavior over anotherChosen/rejected response pairs or reward signals
Task fine-tuningOptimize for a narrow taskLabeled examples
DistillationTransfer behavior from a teacher modelTeacher outputs, logits, or generated traces

Supervised fine-tuning, or SFT, trains against known target outputs. The model is penalized when its predicted tokens differ from the desired response.

Preference alignment goes beyond imitation. Methods such as RLHF, RLAIF, and DPO use human or AI preferences to make some responses more likely than others. This can shape helpfulness, safety, style, verbosity, refusal behavior, and reasoning strategy.

A model named Instruct, Chat, or Assistant has usually gone through SFT and possibly one or more preference stages. The exact recipe should be in the model card; the name alone is not a technical specification.

Full fine-tuning versus parameter-efficient fine-tuning

In a full fine-tune, the training process can update all model weights. This provides maximum freedom but requires substantial memory for gradients, optimizer states, activations, and often higher-precision master weights.

Parameter-efficient fine-tuning, or PEFT, freezes most of the base model and trains a much smaller set of parameters. LoRA is the most common example, but it is not the only PEFT method.

ApproachWhat is trainedStorage for each variantMain tradeoff
Full fine-tuneMost or all base weightsAnother full checkpointMaximum flexibility, highest training cost
LoRASmall low-rank update matricesA compact adapterEfficient and portable, dependent on the base
Prompt / prefix tuningLearned virtual prompt representationsVery small adapterCheap, but less flexible for some tasks
Partial or layer tuningSelected layers or parametersBetween adapter and full checkpointMore capacity with more cost

PEFT reduces trainable parameters. It does not make the frozen base disappear. I still need the base model in memory during training and inference.

What a LoRA actually is

LoRA stands for Low-Rank Adaptation. Instead of updating a large weight matrix directly, LoRA learns two much smaller matrices whose product represents an update to that weight.

In simplified form:

effective weight = base weight + scaled LoRA update
W' = W + scale × B × A

The base matrix W stays frozen. Only A and B are trained.

The word rank describes the inner dimension of those smaller matrices. A higher rank gives the adapter more capacity but increases training cost and adapter size. lora_alpha controls the update's scaling, and target_modules determines which parts of the network receive LoRA layers.

This is why two LoRAs for the same model can differ dramatically. One may target only attention projections at rank 8. Another may target every linear layer at rank 64. Both are “LoRAs,” but they do not have the same capacity, size, or effect.

LoRA works well because many useful model adaptations can be represented with a much lower-rank update than the full weight matrix. The original LoRA paper freezes the pretrained weights and injects trainable rank-decomposition matrices, reducing the number of parameters that need to be trained.

A LoRA is normally not a complete model

A LoRA adapter might contain:

adapter_model.safetensors
adapter_config.json

The Safetensors file contains the learned adapter tensors. The configuration records information such as the PEFT method, rank, scaling, target modules, and often the expected base model.

To use it, I need:

  1. The correct base model architecture
  2. The correct base checkpoint or a compatible revision
  3. A compatible tokenizer and configuration
  4. A runtime that knows how to apply that adapter

An adapter trained for one 27B model cannot be assumed to work on another 27B model. Matching parameter count is not enough. Tensor names, shapes, layer counts, architecture revisions, and tokenizer choices must agree.

This matters just as much for image-generation LoRAs. A LoRA trained for one diffusion-model family is not automatically compatible with another merely because both generate images.

LoRA at runtime versus a merged model

There are two common ways to deploy a LoRA.

Load the base and adapter separately

The runtime loads the base model, loads the adapter, and applies the LoRA update during inference.

This is useful when I want to:

  • Switch among several adapters
  • Combine adapters supported by the runtime
  • Keep one shared base model on disk
  • Preserve the original adapter for more training

The cost is added runtime complexity and, depending on the engine and targeted layers, some inference overhead.

Merge the adapter into the base

Merging calculates the effective weights and saves a new standalone checkpoint. The resulting model no longer needs the LoRA files at inference time.

This is useful when I want to:

  • Convert the result into GGUF
  • Quantize one final deployment model
  • Use a runtime without PEFT support
  • Eliminate adapter application overhead

Merging gives up easy adapter switching and usually creates another full-size checkpoint. It can also be lossy or unsupported when performed directly against aggressively quantized weights. My safest deployment path is to merge into a high-precision base, verify the result, and then quantize the merged model.

Some runtimes can apply a LoRA directly to a quantized GGUF base. That is convenient for experimentation, but it is a different workflow from creating one permanently merged and quantized file.

QLoRA is a training technique, not a special model family

QLoRA combines a frozen quantized base model with trainable LoRA adapters. The original QLoRA work loads the base in 4-bit NormalFloat, performs computation at a higher precision, and backpropagates through the frozen quantized model into the LoRA weights.

The key idea is:

4-bit frozen base + trainable LoRA adapters = much lower fine-tuning memory

QLoRA does not mean the adapter itself is a Q4 model. The adapter weights are normally stored separately in a higher-precision tensor format. It also does not mean the final merged model must remain 4-bit. I can merge the learned update into a higher-precision base checkpoint and then choose a deployment precision.

QLoRA makes fine-tuning far more accessible, but training still needs memory for activations, gradients for the adapter, optimizer state, batches, and sequence length. Inference memory and fine-tuning memory are not interchangeable.

Safetensors: a safe box for tensors

Safetensors is a storage format for tensors—the multidimensional arrays that hold model weights and other learned values. Hugging Face describes it as a safe alternative to pickle with fast, zero-copy loading.

The security distinction matters. Python pickle can reconstruct arbitrary Python objects and may execute code during deserialization. Safetensors stores a restricted header plus raw tensor data; the file is not a general Python-object program.

That does not make every repository containing Safetensors automatically safe. A model repo may include custom Python code, and loading it with options such as trust_remote_code=True can execute that code. Safetensors makes the weight file safer; it does not bless every accompanying file or runtime action.

Safetensors is especially useful in training and GPU inference ecosystems because it:

  • Preserves named tensors and their data types
  • Supports memory mapping and efficient slices
  • Avoids pickle's arbitrary-code deserialization
  • Works well with sharded checkpoints
  • Integrates with Transformers, Diffusers, PEFT, MLX, and other tools

One Safetensors file may not be a complete model

Large checkpoints are often split into files such as:

model-00001-of-00008.safetensors
model-00002-of-00008.safetensors
...
model.safetensors.index.json

The index maps tensor names to shards. The repository also needs configuration and tokenizer files. Copying only one shard will not produce a smaller working model; it produces an incomplete checkpoint.

A typical Transformers-style repository may look like:

config.json
generation_config.json
tokenizer.json
tokenizer_config.json
special_tokens_map.json
chat_template.jinja
model-00001-of-00008.safetensors
...
model.safetensors.index.json

The weights are in Safetensors. The model definition and behavior depend on the rest of the package.

GGUF: an inference package, not a synonym for Q4

GGUF is a binary format from the GGML ecosystem designed to package tensors and model metadata for efficient inference. It is most closely associated with llama.cpp and the many applications built around it.

A GGUF can include:

  • Architecture metadata
  • Model tensors and their data types
  • Tokenizer vocabulary and token metadata
  • Context and positional-encoding parameters
  • Quantization information
  • A chat template and other metadata

The format is designed to be self-describing enough for a compatible runtime to map and execute the model efficiently. It supports memory mapping, which lets the operating system map file regions instead of copying the entire file through a traditional read path.

GGUF is popular for local inference because llama.cpp can run across CPUs and multiple accelerator backends, split layers between CPU and GPU, and support a wide range of block quantizations.

But GGUF does not mean 4-bit. A GGUF can contain F32, F16, BF16, Q8, Q6, Q5, Q4, Q3, Q2, or a mixture of tensor types. The extension tells me the container. The quant name and file size tell me the weight representation.

GGUF also does not guarantee universal compatibility. A runtime must implement the specific model architecture, tensor types, tokenizer behavior, and metadata version. A brand-new model may have valid GGUF files before my installed runtime knows how to execute them.

Safetensors versus GGUF

They overlap, but they are optimized for different parts of the workflow.

QuestionSafetensorsGGUF
Primary roleSafe tensor storage for training and framework-native checkpointsSelf-describing package for efficient local inference
Common ecosystemTransformers, PyTorch, Diffusers, PEFT, MLXllama.cpp, LM Studio, KoboldCpp, compatible local apps
Typical contentsNamed tensors; companion files provide config and tokenizerTensors plus substantial model and tokenizer metadata
QuantizationCan store original or quantized tensors supported by the frameworkCommonly used with GGML block quantizations; can also store higher precision
Fine-tuningNatural source format for training and adaptersPrimarily an inference format, though adapters can also be converted
ShardingCommon for large modelsOften one file, but large GGUF models can be split
Hardware pathDepends on PyTorch/framework kernelsDesigned for portable CPU/GPU inference backends
Best choiceTraining, fine-tuning, framework-native GPU servingSimple local deployment and flexible hardware offload

Neither format is inherently smarter or higher quality. If both contain equivalent weights at equivalent precision and both runtimes implement the model correctly, the model has the same learned capability. Differences come from quantization, kernels, prompt formatting, sampler settings, and implementation details—not the filename's personality.

Converting Safetensors to GGUF

The normal llama.cpp deployment path is:

Transformers checkpoint
        ↓
BF16/F16 GGUF conversion
        ↓
GGUF quantization
        ↓
llama.cpp-compatible runtime

The converter reads the architecture configuration, maps tensor names and layouts, imports tokenizer data, writes GGUF metadata, and serializes the weights. A quantizer can then create smaller variants from the high-precision GGUF.

This is why I prefer to keep the original high-precision checkpoint or at least one high-quality intermediate. Re-quantizing an already low-bit model compounds approximation error. A Q4 file is a deployment artifact, not the ideal master copy for producing every future format.

Conversion can fail or silently behave badly when:

  • The architecture is not supported yet
  • Tensor layouts changed in a new model revision
  • The tokenizer or special tokens are missing
  • A multimodal projector was not converted
  • The chat template is wrong or absent
  • The model uses an unsupported quantization or custom layer

“Converted successfully” only means the file was written. I still test generation, special tokens, tool calling, long context, and task quality.

Other model formats I run into

Safetensors and GGUF cover much of my local workflow, but they are not the only formats.

FormatWhere it fits
.bin, .pt, .pthPyTorch checkpoints; may use pickle and require more trust when loading
ONNXPortable computation graph for inference across ONNX runtimes
EXL2GPU-focused quantization format used by ExLlamaV2-style runtimes
GPTQ / AWQ checkpointsQuantized weight layouts used by supported GPU inference engines
MLXApple-silicon-focused model packages for the MLX ecosystem
TensorRT enginesHighly optimized, hardware- and configuration-specific NVIDIA inference artifacts
Adapter SafetensorsSmall PEFT or LoRA updates that require a compatible base

Some names in that table describe a file format, some a quantization method, and some an execution ecosystem. The industry does not maintain a clean taxonomy, so I always ask two questions: What is stored? Which runtime is expected to load it?

Model types by architecture and task

The training and file-format labels sit on top of another distinction: what kind of network and output the model was designed to produce.

Decoder-only causal language models

These predict the next token from previous tokens. Most modern chat, coding, and agentic LLMs belong here. They are naturally suited to open-ended generation.

Encoder-only models

These build representations from the full input rather than generating long continuations one token at a time. They are common for classification, named-entity recognition, embeddings, and some rerankers.

Encoder-decoder models

The encoder reads an input sequence and the decoder generates a target sequence. Translation and summarization are classic uses.

Embedding models

An embedding model turns text, images, or other inputs into vectors. Similar items should land near one another in vector space. These models power semantic search and retrieval; they are not chat models even when both use Transformer components.

Rerankers

A reranker scores a query-document pair more precisely than a basic embedding similarity search. A common retrieval pipeline uses embeddings to find candidates quickly and a reranker to reorder the best few.

Vision-language and multimodal models

A VLM combines a language model with a vision encoder, projector, or integrated multimodal architecture. The text weights alone may load while image understanding remains unavailable because a required projector or vision tower is missing.

Diffusion and other generative-media models

Image, audio, and video generators may use Transformers, diffusion components, VAEs, text encoders, and LoRAs. They can use Safetensors or GGUF too, but their pipeline and runtime requirements differ from an autoregressive LLM.

The extension does not define the task. A .safetensors file could be an LLM shard, a LoRA, a diffusion transformer, a VAE, a text encoder, or an embedding model.

Base, instruct, reasoning, coding, and agentic are training labels

Model names often include behavior-oriented labels:

LabelWhat I expectWhat I still verify
BaseGeneral pretrained completion modelTraining data, license, context
Instruct / ChatRequest following and conversationChat template and alignment quality
ReasoningMore deliberate multi-step behaviorToken cost, mode controls, benchmark relevance
CoderMore code and software-task trainingLanguages, fill-in-the-middle, tool use
AgenticTool use and multi-step environment interactionHarness compatibility and recovery behavior
EmbeddingVector output for retrieval or similarityVector size, pooling, supported languages
RerankerQuery-document relevance scoringInput length and serving cost
Abliterated / uncensoredAlignment or refusal behavior was modifiedCapability loss, provenance, and safety implications

These labels are useful search terms, not guarantees. A coder model can be poor at editing a real repository. An agentic model can produce perfect tool-call syntax but fail to recover from a test error. I treat the model card and my own evaluations as the specification.

How to read a model repository name

Suppose I see:

Example-27B-Instruct-LoRA
Example-27B-Instruct-AWQ
Example-27B-Instruct-GGUF
Example-27B-Instruct-Q4_K_M.gguf

I read it from left to right:

  1. Example identifies the model family or architecture.
  2. 27B gives the approximate parameter count.
  3. Instruct describes the training stage or behavior.
  4. LoRA, AWQ, or GGUF describes the artifact, quantization, or ecosystem.
  5. Q4_K_M identifies the specific quantized tensor recipe.

The LoRA repository probably does not contain full weights. The AWQ repository likely contains a complete quantized checkpoint for an AWQ-compatible loader. The GGUF repository may contain several downloadable quant levels. The final filename identifies one of those files.

Community naming is inconsistent, so I verify the contents instead of trusting the suffix.

The files I check before downloading

The largest file gets the attention, but these are just as important:

Model card

The model card should identify the base model, training recipe, intended uses, limitations, evaluation results, prompt format, and license. If a fine-tune does not clearly name its base, I consider that a deployment risk.

License

Open weights do not automatically mean unrestricted use. The base model, fine-tune dataset, adapter, and redistributed merge may have different conditions. Merging weights does not erase their licenses.

Configuration

config.json describes the architecture. An unsupported model_type or custom layer may require a newer runtime or repository code.

Tokenizer

Tokenizer files define how text becomes token IDs. A mismatch can corrupt output even if every model tensor loads correctly.

Chat template

The template formats system, user, assistant, and tool messages into the token pattern used during instruction training. The wrong template can cause broken role handling, repeated text, poor tool calls, or immediate end-of-sequence tokens.

Checksums and provenance

I prefer files traceable to the model author or a reputable converter. Exact hashes matter when a model is updated without a clear filename change.

My preferred workflows

I only want to run the model locally

I choose an instruct or specialized checkpoint, find a GGUF quant that fits with context headroom, confirm my runtime supports the architecture, and use the model's expected chat template.

I want maximum GPU serving performance

I start with the original Safetensors checkpoint and choose the quantization and engine supported by my GPU stack—BF16, FP8, AWQ, GPTQ, or another optimized path. GGUF is wonderfully portable, but it is not automatically the fastest format for every NVIDIA deployment.

I want to create a custom personality or domain behavior

I keep the original base or instruct checkpoint in Safetensors, train a LoRA or QLoRA adapter, evaluate the adapter against the exact base, and preserve both before merging.

I want one deployable GGUF from my LoRA

My clean path is:

exact base checkpoint
        +
LoRA adapter
        ↓
merge at high precision
        ↓
export or convert to high-precision GGUF
        ↓
quantize the merged GGUF
        ↓
test prompt format and real tasks

That workflow gives me a reusable adapter, a high-quality merged master, and one or more deployment quants. It costs more disk space, but it avoids turning the smallest artifact into the only copy I can build from.

My actual rule

I do not ask, “Is GGUF better than Safetensors?” They solve different problems.

I ask:

  1. What architecture and task is this model built for?
  2. Is this a base, instruct, aligned, or specialized checkpoint?
  3. Is the download a complete model or only an adapter?
  4. What precision are the weights stored in?
  5. Which runtime is expected to load them?
  6. Are the tokenizer, template, license, and base-model revision clear?

Safetensors is where I prefer to keep trainable, framework-native source checkpoints and adapters. GGUF is where I prefer to package a practical local-inference build. LoRA is how I avoid duplicating an entire base model for every experiment. Merging is how I turn one successful experiment into a standalone deployment artifact.

The file is not the model. It is one layer of the model's history.

Sources