Understanding Its Purpose First
Next-token prediction estimates the likelihood of various tokens following given content.
Understanding with an Example
When given the input "the capital of France is," many continuations in the token vocabulary receive different scores. The process by which the model ultimately generates tokens related to "Paris" is based on numerical computation, not a lookup in an internal fact table row by row.
The capital of France is...
Calculate probabilities for candidate tokens
Compare with target during training; select continuation during use
Diving Deeper
During training, the existing text provides the actual subsequent tokens, and the program increases the prediction probability of these target tokens. During use, there is no pre-given answer, and the generated tokens are added to the subsequent conditions. A Chinese character, word, or punctuation mark does not have a fixed one-to-one correspondence with a token; the specific mapping depends on the tokenizer.
What It Cannot Tell You
"Predicting continuations" describes the generation mechanism and does not mean the capability is limited to completing a few characters; nor does it guarantee the accuracy of answers based on this. The text in the diagram is for teaching purposes and does not mimic the output of any actual tokenizer.
What to Learn Next
Sources
- Hugging Face · How do Transformers work?: Pretraining and fine-tuning, different architectures and training objectives.
- Hugging Face · Caching: Reusing historical K/V across layers in autoregressive inference; attention Q/K/V and cache memory overhead.
Information verified on 2026-09-09; original papers are used to illustrate mechanisms, and examples in the text are for teaching purposes only.