Understanding Its Purpose First

SGD uses sampled examples to estimate the gradient, then updates parameters in the direction that reduces the loss.

Understanding Through an Example

Suppose you have one million training examples—you don't need to review all of them before making an adjustment. Training typically uses a small batch of examples to compute the gradient and update, so each step has fluctuation, but computation becomes more feasible.

Visual explanationSee the Process Clearly
01Sample a batch

Estimate current gradient

02Take a step along negative gradient

Learning rate controls step size

03Continue with New Samples

Multi-round Cumulative Improvement

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going One Level Deeper

Strictly speaking, single-example updates are stochastic gradient descent, while mini-batch updates are mini-batch SGD; in practice, both are commonly referred to as SGD. The basic rule is: "new parameters = old parameters − learning rate × gradient." Implementations may also incorporate momentum to moderate frequent direction changes.

What It Cannot Explain

Randomness doesn't mean arbitrarily changing parameters—the gradient still comes from the loss and the examples. Mini-batch gradients are merely an estimate; improvement in one batch doesn't guarantee improvement across other examples. This rule serves as a starting point for understanding training, not as an assertion about all large model training recipes.

What to Explore Next

References

Information verified on 2026-09-09; original papers are cited to explain mechanisms, and the examples in the text are for instructional purposes only.