Pretraining vs Fine-Tuning: How Do LLM Stages Differ?
Pretraining develops broad capabilities from a large corpus. Supervised fine-tuning teaches a model to produce desired responses from examples. Alignment methods train it toward preferences or verifiable rewards. These describe different training objectives; a useful model can pass through all three.
Start with the behavior that needs to change. A missing document, a repeated classification error, and an unsafe action require different interventions.
What changes at each stage?
| Stage | Training signal | Result to evaluate |
|---|---|---|
| Pretraining | For an autoregressive LLM, predict the next token in corpus text | General language and task capabilities |
| Supervised fine-tuning, or SFT | Input and desired-response examples | Behavior on the demonstrated task |
| Preference or reward training | Preferred response pairs, learned scores, or rule-based checks | The rewarded behavior and unwanted side effects |
The Llama 3 technical report distinguishes large-scale pretraining from post-training that develops instruction following and other capabilities. Pretraining is not always next-token prediction: encoder and denoising models can use other objectives.
SFT is usually the more direct starting point when reviewed examples show the output you want. For example, an extraction dataset can pair a document with the correct fields. Preference training fits comparisons such as which of two answers better follows a policy. DPO optimizes preference pairs directly, while a conventional RLHF pipeline trains a reward model and then updates the policy through reinforcement learning.
Choose from the observed failure
If answers need changing or private facts, first supply those facts through retrieval. If the model repeatedly uses the wrong labels despite a clear prompt, test SFT on reviewed examples. If several acceptable responses differ in a preference you can judge consistently, consider preference training.
Keep a held-out evaluation set across each change. Check the target behavior, general task regressions, refusals, and cost. A lower training loss does not establish that the deployed system improved.
For the application choice, see fine-tuning vs RAG vs prompting. The training-stages section of the LLM Engineering Guide provides the wider context.