Understanding Its Purpose
A loss function converts the gap between predictions and training targets into a single number for the training program to optimize.
Understanding Through an Example
In the training text, the target token is "textbook". Assigning it a 10% probability typically incurs a larger loss than assigning it a 60% probability; this way, the program knows which adjustment direction better aligns with the training target.
Target tokens currently only 10%
Correct continuation given by training samples
More deviant from target, typically greater penalty
Going One Level Deeper
Common language modeling objectives use cross-entropy. For a single known target token, this can first be understood as the negative log probability: L = −log p. p represents the probability the model assigns to the target token; for instance, under the natural logarithm, p=0.1 yields approximately 2.30, while p=0.6 yields approximately 0.51. Actual training aggregates across many positions and samples.
What It Cannot Tell You
A smaller number only indicates better alignment with this mathematical objective; it does not directly indicate better aesthetics, greater suitability for the classroom, or that all facts are correct. Whether the task is genuinely good still requires separate evaluation.
Where to Go from Here
Sources
- PyTorch · Optimizing Model Parameters: Batches, learning rates, loss, gradients, SGD updates, and test set evaluation.
Materials verified on 2026-09-09; original papers are used to illustrate mechanisms, and the examples in the text are instructional demonstrations.