Understanding Its Role First
The optimizer combines gradients with update rules to determine exactly how the parameters should be adjusted in each step.
Understanding Through an Example
Backpropagation provides information about "which direction to adjust," while the optimizer decides what rule to use for taking a step. SGD and Adam both operate at this level; the Transformer operates at the model architecture level.
Derived from current step loss
Combines learning rate with historical state
Used for next iteration computation
Going a Layer Deeper
The learning rate controls the scale of updates; the batch determines which samples are used per step; some optimizers also maintain statistics of past gradients. SGD can serve as a foundational understanding, while Adam adapts per-parameter updates based on first and second-order moment estimates of the gradients. They are not sequential course steps that must be followed in order.
What It Does Not Explain
Choosing an optimizer does not replace the need for thoughtful data, objectives, and evaluation design. Training methods like PPO still rely on gradient-based optimizers internally to update parameters—these operate at different levels of abstraction.
Next Steps
- Stochastic Gradient Descent (SGD)
- Adaptive Moment Estimation (Adam)
- Gradient
- Proximal Policy Optimization (PPO)
References
- PyTorch · Optimizing Model Parameters: Batch, learning rate, loss, gradients, SGD updates, and test set evaluation.
- Adam: A Method for Stochastic Optimization: Section 2 and Algorithm 1: first and second-order moments, bias correction, parameter updates.
References verified on 2026-09-09; the original paper is used to illustrate mechanisms, and the examples in the text are for instructional purposes.