Understanding Its Function
Attention allows one position to gather information from other permitted positions according to computed weights.
Understanding Through an Example
「Xiao Lin handed the book to the teacher, and the teacher opened it.」 Understanding the final 「it」 requires integrating the preceding context. The bold connections in the diagram are for explaining information aggregation only, not measurements of real model internal weights.
“它” 汇集前文信息 · 强调联系仅为教学示意
What information is needed for this step
Calculate matching scores with each position
Aggregate by normalized weighted sum
Going a Step Deeper
Common scaled dot-product attention computes similarity scores using query Q and key K, then performs weighted summation over values V after normalization. These are vectors obtained through value representation transformations, not three text files. Multiple attention heads can interact in different subspaces.
What It Cannot Explain
Attention weights do not represent human psychological focus, nor can they alone prove the model's causal reasoning process. Causal language models cannot peek at future answers when generating the current position; the diagram uses content that has already appeared.
Next Steps
References
- Attention Is All You Need: Transformer is a model architecture centered on attention; the original paper includes an encoder and decoder.
- Hugging Face · Caching: Reuse of historical K/V across layers during autoregressive inference; attention Q/K/V and cached memory overhead.
Data verified on 2026-09-09; the original paper is used to explain the mechanism, and the example in the text is a teaching illustration.