Understanding Its Role

The gradient describes how the loss will change when parameters are adjusted slightly in each direction at the current position.

Understanding Through an Example

Consider a simple curve: loss L = (w − 3)², with the current parameter w = 1. Here the gradient is −4; after subtracting the gradient multiplied by the learning rate in standard gradient descent, w moves closer to 3.

Visual explanationConsider one parameter to understand one update step
当前位置走一小步损失参数
01Current position

w = 1, loss = 4

02Local direction

gradient = −4

03Small update

Learning rate 0.1 → w = 1.4

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going Deeper

The gradient points in the direction of steepest local increase, so standard gradient descent moves in the opposite direction. After the update in the example, the loss becomes 2.56, which is less than 4. Real models have many parameters, and the gradient contains many components; practical optimizers may also use historical statistics to adjust the step size.

What It Doesn't Tell You

This single-parameter example only illustrates local changes. It does not guarantee that every step of actual training will decrease the loss, nor does it guarantee finding the global optimum. A learning rate that is too large may skip over minima, and changes in data batches also introduce fluctuations.

Where to Go Next

References

References verified on 2026-09-09; original papers are used to explain mechanisms, and examples in the text are for instructional purposes.