Understanding Its Purpose

A loss function converts the gap between predictions and training targets into a single number for the training program to optimize.

Understanding Through an Example

In the training text, the target token is "textbook". Assigning it a 10% probability typically incurs a larger loss than assigning it a 60% probability; this way, the program knows which adjustment direction better aligns with the training target.

Visual explanationSee the Process Clearly
01Predicted probability

Target tokens currently only 10%

02Known target

Correct continuation given by training samples

03Loss value

More deviant from target, typically greater penalty

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going One Level Deeper

Common language modeling objectives use cross-entropy. For a single known target token, this can first be understood as the negative log probability: L = −log p. p represents the probability the model assigns to the target token; for instance, under the natural logarithm, p=0.1 yields approximately 2.30, while p=0.6 yields approximately 0.51. Actual training aggregates across many positions and samples.

What It Cannot Tell You

A smaller number only indicates better alignment with this mathematical objective; it does not directly indicate better aesthetics, greater suitability for the classroom, or that all facts are correct. Whether the task is genuinely good still requires separate evaluation.

Where to Go from Here

Sources

Materials verified on 2026-09-09; original papers are used to illustrate mechanisms, and the examples in the text are instructional demonstrations.