Understanding Its Purpose
GRPO generates a group of answers to the same problem and uses the relative differences in rewards within the group to guide policy updates.
Learning Through an Example
For example, four answers are generated for the same math problem, yielding rewards of 0, 1, 1, and 0 upon checking. Training compares the relative performance within this group; these numbers are teaching examples, and actual tasks require well-designed rewards.
Sample a set of candidates
Quality relative to average reward
Combine with clipping and other objective mechanisms
A Step Deeper
The original DeepSeekMath method uses within-group rewards to estimate a baseline, thereby eliminating the need for a separate value model commonly used in PPO. In the outcome supervision setting, the reward is subtracted by the group mean, then normalized by the group standard deviation to compute the advantage; process-level supervision can also be used.
What It Does Not Imply
Rewards can come from the model or from verifiable rules—no need for all of them to come from humans. When all members in the group receive the same score, the relative signal is limited, and real implementations need to handle cases such as zero variance. GRPO is not about selecting only the highest-scoring answer to show to users, nor is it used by every model.
Where to Go Next
Source Material
- DeepSeekMath · GRPO: Section 4.1: Candidate groups for the same problem, relative rewards, free from a separate value model; outcome supervision and process supervision.
Material verified on 2026-09-09; the original paper is used to illustrate the mechanism, and the numbers in the example are teaching illustrations.