Understanding Its Role
Training data is the sample data that models use for learning; the coverage and quality of the data affect what the model learns.
Understanding Through an Example
If all lesson preparation examples come from high school materials, even with large quantities, the model may still fail to learn how to organize activities for elementary school students. If test questions have already been mixed into the training samples, high test scores may exaggerate actual capability.
Used to update parameters
Help tune settings and select approaches
Reserved for independent evaluation
Going Deeper
Data processing typically involves filtering, cleaning, deduplication, and organizing samples. Different training tasks require different formats: language pretraining can construct next-token objectives from text, supervised fine-tuning often uses instructions and example answers, and preference training uses comparative information between candidate answers.
What It Cannot Tell You
The three datasets cannot simply have different names while containing largely repeated content. Actual splits need to consider same sources, similar question types, and time ranges; otherwise, evaluation may still leak. The three splits here are simplified for teaching purposes.
Continue Learning About
References
- PyTorch · Optimizing Model Parameters: Batches, learning rate, loss, gradients, SGD updates, and test set evaluation.
- Training language models to follow instructions with human feedback: InstructGPT examples: demonstration data, human rankings, and post-training from feedback.
References verified on 2026-09-09; original papers are used to explain mechanisms, and examples in the text are for instructional purposes.