Understanding Its Role
Supervised fine-tuning uses demonstrations of inputs paired with ideal outputs to help an existing model learn response patterns that better align with task requirements.
Understanding Through an Example
A sample might consist of: "Design a 10-minute activity for 7th-grade students," along with an appropriate answer prepared by a teacher. The model learns the task structure, expression style, and content patterns from the demonstrations.
Grade, Objective and Time
Desired Results
Increase probability of demonstration outputs
Going a Step Deeper
In language models, training is commonly performed by predicting tokens in the demonstration answers. Sample design matters significantly: instruction diversity, demonstration quality, data coverage, and evaluation all influence outcomes. SFT is a training approach that uses supervised demonstrations. It can update a large number of parameters or be combined with LoRA to update only a small set of adapter parameters.
What It Does Not Mean
When a regular user pastes a few examples in a single conversation, this is typically called in-context demonstration and does not constitute SFT. Both approaches can influence responses, but whether parameters are updated is the key distinction.
Where to Go Next
References
- Training language models to follow instructions with human feedback: InstructGPT example: demonstration data, human ranking, post-training via feedback.
- LoRA: Low-Rank Adaptation of Large Language Models: Freezes pretrained weights and trains low-rank matrices; a parameter-efficient adaptation method.
References verified on 2026-09-09; original papers are cited to explain mechanisms, and examples in the text are for instructional purposes only.