RLHF vs DPO vs GRPO: How Do Alignment Methods Differ?

Conventional PPO-based RLHF trains a policy against a reward model learned from human preferences. DPO learns directly from preferred and rejected response pairs. GRPO samples groups of responses and compares their rewards to update the policy without PPO’s learned critic.

RLHF describes the human-feedback approach; PPO, DPO, and GRPO describe training methods. GRPO can use verifiable rules or a learned reward model rather than human judgments.

Choose by the feedback you can trust

MethodTraining feedbackFresh responses during training?Useful condition
PPO-based RLHFLearned reward scoreYesOnline exploration is needed and the reward model is tested
Standard DPOPreferred/rejected pairsNo; fixed pairsReviewed comparisons cover the behavior you want
GRPORewards relative to other responses in the groupYesThe task has informative, repeatable reward checks

The DPO paper derives a preference-pair objective that avoids the separate reward-model training and online reinforcement-learning loop of conventional RLHF. That simplifies implementation, but fixed pairs cannot explore responses absent from the dataset.

DeepSeekMath introduced GRPO’s critic-free group baseline. For a math task, several sampled answers can receive correctness rewards. If every answer receives the same reward, the group supplies little information about which behavior to improve.

Count the actual training state

A conventional PPO pipeline has policy, reference, reward, and value-model roles. They need not be four full copies resident on each GPU. Sharing, adapters, sharding, and offload change memory use.

GRPO’s reference model is also implementation-dependent. The current TRL GRPO trainer defaults to a zero KL coefficient and does not load a reference model in that configuration. Enabling reference-based KL regularization or a learned reward model adds state.

Evaluate the trained model on held-out tasks and inspect reward exploitation. A passing unit test, a preference score, or a correct final number checks only what the evaluator measures. Compare general-task regressions and generation cost as well as the rewarded metric.

The alignment-methods section of the LLM Engineering Guide gives the wider training context.