Understanding Its Purpose First
SGD uses sampled examples to estimate the gradient, then updates parameters in the direction that reduces the loss.
Understanding Through an Example
Suppose you have one million training examples—you don't need to review all of them before making an adjustment. Training typically uses a small batch of examples to compute the gradient and update, so each step has fluctuation, but computation becomes more feasible.
Estimate current gradient
Learning rate controls step size
Multi-round Cumulative Improvement
Going One Level Deeper
Strictly speaking, single-example updates are stochastic gradient descent, while mini-batch updates are mini-batch SGD; in practice, both are commonly referred to as SGD. The basic rule is: "new parameters = old parameters − learning rate × gradient." Implementations may also incorporate momentum to moderate frequent direction changes.
What It Cannot Explain
Randomness doesn't mean arbitrarily changing parameters—the gradient still comes from the loss and the examples. Mini-batch gradients are merely an estimate; improvement in one batch doesn't guarantee improvement across other examples. This rule serves as a starting point for understanding training, not as an assertion about all large model training recipes.
What to Explore Next
References
- PyTorch · Optimizing Model Parameters: batch, learning rate, loss, gradient, SGD update, and test set evaluation.
- Adam: A Method for Stochastic Optimization: Section 2 and Algorithm 1: first and second moment estimates, bias correction, parameter update.
Information verified on 2026-09-09; original papers are cited to explain mechanisms, and the examples in the text are for instructional purposes only.