Understanding Its Role First

The optimizer combines gradients with update rules to determine exactly how the parameters should be adjusted in each step.

Understanding Through an Example

Backpropagation provides information about "which direction to adjust," while the optimizer decides what rule to use for taking a step. SGD and Adam both operate at this level; the Transformer operates at the model architecture level.

Visual explanationSee the Process Clearly
01Input gradient

Derived from current step loss

02Update rule

Combines learning rate with historical state

03New parameters

Used for next iteration computation

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going a Layer Deeper

The learning rate controls the scale of updates; the batch determines which samples are used per step; some optimizers also maintain statistics of past gradients. SGD can serve as a foundational understanding, while Adam adapts per-parameter updates based on first and second-order moment estimates of the gradients. They are not sequential course steps that must be followed in order.

What It Does Not Explain

Choosing an optimizer does not replace the need for thoughtful data, objectives, and evaluation design. Training methods like PPO still rely on gradient-based optimizers internally to update parameters—these operate at different levels of abstraction.

Next Steps

References

References verified on 2026-09-09; the original paper is used to illustrate mechanisms, and the examples in the text are for instructional purposes.