Understanding Its Purpose
Pretraining allows models to first acquire foundational capabilities from large datasets that can be applied to subsequent tasks.
Understanding Through an Example
In language models, large amounts of written text, code, and other documents provide patterns of language. The model first learns broad continuation abilities, and only then might it further learn to produce a well-structured lesson plan according to specific requirements.
Content and quality affect the learned patterns
For example, next token prediction
Obtain capabilities for further adaptation
Going Deeper
For common autoregressive language models, training signals can be generated from the text itself: the preceding context serves as the condition, and the actual subsequent tokens serve as the target. Not all pretraining predicts the next token—other architectures may employ objectives like mask recovery. Pretraining is a stage name, not a single unique algorithm.
What It Doesn't Tell You
A base model is not a complete application. It does not automatically gain file access, PowerPoint formatting tools, or fact-verification capabilities. General users interact with systems that have undergone further training and are integrated with application features.
What to Explore Next
Sources
- Hugging Face · How do Transformers work?: Pretraining vs. fine-tuning, different architectures and training objectives.
- Attention Is All You Need: The Transformer is an attention-centric model architecture; the original paper covers encoders and decoders.
Facts verified on 2026-09-09; the original paper is cited to illustrate mechanisms, and the examples in the text are for instructional purposes.