Understanding Its Role
The gradient describes how the loss will change when parameters are adjusted slightly in each direction at the current position.
Understanding Through an Example
Consider a simple curve: loss L = (w − 3)², with the current parameter w = 1. Here the gradient is −4; after subtracting the gradient multiplied by the learning rate in standard gradient descent, w moves closer to 3.
w = 1, loss = 4
gradient = −4
Learning rate 0.1 → w = 1.4
Going Deeper
The gradient points in the direction of steepest local increase, so standard gradient descent moves in the opposite direction. After the update in the example, the loss becomes 2.56, which is less than 4. Real models have many parameters, and the gradient contains many components; practical optimizers may also use historical statistics to adjust the step size.
What It Doesn't Tell You
This single-parameter example only illustrates local changes. It does not guarantee that every step of actual training will decrease the loss, nor does it guarantee finding the global optimum. A learning rate that is too large may skip over minima, and changes in data batches also introduce fluctuations.
Where to Go Next
References
- PyTorch · Optimizing Model Parameters: Batches, learning rate, loss, gradients, SGD updates, and test set evaluation.
- PyTorch · Automatic Differentiation: Backpropagation computes parameter gradients via the chain rule; it is separate from optimizer updates.
References verified on 2026-09-09; original papers are used to explain mechanisms, and examples in the text are for instructional purposes.