Understanding Its Role

Training data is the sample data that models use for learning; the coverage and quality of the data affect what the model learns.

Understanding Through an Example

If all lesson preparation examples come from high school materials, even with large quantities, the model may still fail to learn how to organize activities for elementary school students. If test questions have already been mixed into the training samples, high test scores may exaggerate actual capability.

Visual explanationView different effects together
01Training set

Used to update parameters

02Validation set

Help tune settings and select approaches

03Test set

Reserved for independent evaluation

Each item is a different role or optional method, not meaning they must be executed in order.

Going Deeper

Data processing typically involves filtering, cleaning, deduplication, and organizing samples. Different training tasks require different formats: language pretraining can construct next-token objectives from text, supervised fine-tuning often uses instructions and example answers, and preference training uses comparative information between candidate answers.

What It Cannot Tell You

The three datasets cannot simply have different names while containing largely repeated content. Actual splits need to consider same sources, similar question types, and time ranges; otherwise, evaluation may still leak. The three splits here are simplified for teaching purposes.

Continue Learning About

References

References verified on 2026-09-09; original papers are used to explain mechanisms, and examples in the text are for instructional purposes.