Understanding Its Role

Supervised fine-tuning uses demonstrations of inputs paired with ideal outputs to help an existing model learn response patterns that better align with task requirements.

Understanding Through an Example

A sample might consist of: "Design a 10-minute activity for 7th-grade students," along with an appropriate answer prepared by a teacher. The model learns the task structure, expression style, and content patterns from the demonstrations.

Visual explanationSee the Process Clearly
01Task Input

Grade, Objective and Time

02Demonstration Answer

Desired Results

03Continue Training

Increase probability of demonstration outputs

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going a Step Deeper

In language models, training is commonly performed by predicting tokens in the demonstration answers. Sample design matters significantly: instruction diversity, demonstration quality, data coverage, and evaluation all influence outcomes. SFT is a training approach that uses supervised demonstrations. It can update a large number of parameters or be combined with LoRA to update only a small set of adapter parameters.

What It Does Not Mean

When a regular user pastes a few examples in a single conversation, this is typically called in-context demonstration and does not constitute SFT. Both approaches can influence responses, but whether parameters are updated is the key distinction.

Where to Go Next

References

References verified on 2026-09-09; original papers are cited to explain mechanisms, and examples in the text are for instructional purposes only.