Understanding Its Function

Attention allows one position to gather information from other permitted positions according to computed weights.

Understanding Through an Example

「Xiao Lin handed the book to the teacher, and the teacher opened it.」 Understanding the final 「it」 requires integrating the preceding context. The bold connections in the diagram are for explaining information aggregation only, not measurements of real model internal weights.

Visual explanationUnderstanding information aggregation from word relationships
小林书老师翻开它

“它” 汇集前文信息 · 强调联系仅为教学示意

01Current query Q

What information is needed for this step

02Candidate key K

Calculate matching scores with each position

03Content value V

Aggregate by normalized weighted sum

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going a Step Deeper

Common scaled dot-product attention computes similarity scores using query Q and key K, then performs weighted summation over values V after normalization. These are vectors obtained through value representation transformations, not three text files. Multiple attention heads can interact in different subspaces.

What It Cannot Explain

Attention weights do not represent human psychological focus, nor can they alone prove the model's causal reasoning process. Causal language models cannot peek at future answers when generating the current position; the diagram uses content that has already appeared.

Next Steps

References

  • Attention Is All You Need: Transformer is a model architecture centered on attention; the original paper includes an encoder and decoder.
  • Hugging Face · Caching: Reuse of historical K/V across layers during autoregressive inference; attention Q/K/V and cached memory overhead.

Data verified on 2026-09-09; the original paper is used to explain the mechanism, and the example in the text is a teaching illustration.