4 Direct Preference Optimization
Motivated by the challenges of applying reinforcement learning algorithms on large-scale problems such as fine-tuning language models, our goal is to derive a simple approach for policy optimization using preferences directly. Unlike prior RLHF methods, which learn a reward and then optimize it via RL, our approach leverages a particular choice of reward model parameterization that enables extraction of its optimal policy in closed form, without an RL training loop. As we will describe next in detail, our key insight is to leverage an analytical mapping from reward functions to optimal policies, which enables us to transform a loss function over reward functions into a loss function over policies. This change-of-variables approach avoids fitting an explicit, standalone reward model, while still optimizing under existing models of human preferences, such as the Bradley-Terry model. In essence, the policy network represents both the language model and the (implicit) reward.
Deriving the DPO objective
We start with the same RL objective as prior work, Eq. 3, under a general reward function . Following prior work [31], [30], [19], [15], it is straightforward to show that the optimal solution to the KL-constrained reward maximization objective in Eq. 3 takes the form:
where is the partition function. See Appendix A.1 for a complete derivation. Even if we use the MLE estimate of the ground-truth reward function , it is still expensive to estimate the partition function [19], [15], which makes this representation hard to utilize in practice. However, we can rearrange Eq. 4 to express the reward function in terms of its corresponding optimal policy , the reference policy , and the unknown partition function . Specifically, we first take the logarithm of both sides of Eq. 4 and then with some algebra we obtain:
We can apply this reparameterization to the ground-truth reward and corresponding optimal model . Fortunately, the Bradley-Terry model depends only on the difference of rewards between two completions, i.e., . Substituting the reparameterization in Eq. 5 for into the preference model Eq. 1, the partition function cancels, and we can express the human preference probability in terms of only the optimal policy and reference policy . Thus, the optimal RLHF policy under the Bradley-Terry model satisfies the preference model:
The derivation is in Appendix A.2. While Eq. 6 uses the Bradley-Terry model, we can similarly derive expressions under the more general Plackett-Luce models [32], [23], shown in Appendix A.3. Now that we have the probability of human preference data in terms of the optimal policy rather than the reward model, we can formulate a maximum likelihood objective for a parametrized policy . Analogous to the reward modeling approach (i.e. Eq. 2), our policy objective becomes:
This way, we fit an implicit reward using an alternative parameterization, whose optimal policy is simply . Moreover, since our procedure is equivalent to fitting a reparametrized Bradley-Terry model, it enjoys certain theoretical properties, such as consistencies under suitable assumption of the preference data distribution [4]. In Section 5, we further discuss theoretical properties of DPO in relation to other works.
What does the DPO update do?
For a mechanistic understanding of DPO, it is useful to analyze the gradient of the loss function . The gradient with respect to the parameters can be written as:
where is the reward implicitly defined by the language model and reference model (more in Section 5). Intuitively, the gradient of the loss function increases the likelihood of the preferred completions and decreases the likelihood of dispreferred completions . Importantly, the examples are weighed by how much higher the implicit reward model rates the dispreferred completions, scaled by , i.e. how incorrectly the implicit reward model orders the completions, accounting for the strength of the KL constraint. Our experiments suggest the importance of this weighting, as a naïve version of this method without the weighting coefficient can cause the language model to degenerate (Appendix Table 3).
DPO outline
The general DPO pipeline is as follows: 1) Sample completions for every prompt , label with human preferences to construct the offline dataset of preferences and 2) optimize the language model to minimize for the given and and desired . In practice, one would like to reuse preference datasets publicly available, rather than generating samples and gathering human preferences. Since the preference datasets are sampled using , we initialize whenever available. However, when is not available, we initialize by maximizing likelihood of preferred completions , that is,
This procedure helps mitigate the distribution shift between the true reference distribution which is unavailable, and used by DPO. Further details related to the implementation and hyperparameters can be found in Appendix B.