23  Reinforcement Learning with Large Models

Core idea. Reinforcement learning with large models studies how to update a high-dimensional generative policy using human preferences, learned reward models, and regularization. Mathematically, many alignment procedures can be viewed as statistical estimation plus KL-regularized policy optimization:

\[ \max_{\pi}\n\mathbb E_{x\sim \rho,\; y\sim \pi(\cdot\mid x)}[r(x,y)] - \beta\,\mathbb E_{x\sim \rho} \left[D_{\mathrm{KL}}(\pi(\cdot\mid x)\|\pi_0(\cdot\mid x))\right]. \]

Here \(x\) is a prompt, \(y\) is a generated response, \(\pi_0\) is a reference model, \(r\) is a scalar reward or preference score, and \(\beta>0\) controls how far the updated policy may move away from the reference model.

23.1 Learning goals

After reading this chapter, students should be able to:

  1. interpret a language model as a stochastic policy over token sequences;
  2. distinguish contextual-bandit and sequential-MDP views of text generation;
  3. formulate preference learning as a statistical estimation problem;
  4. derive the Bradley-Terry preference model from pairwise comparisons;
  5. explain reward-model training as logistic regression on response pairs;
  6. derive the closed-form optimizer of a KL-regularized policy objective;
  7. explain why KL regularization controls over-optimization and distribution shift;
  8. connect PPO-style RLHF to the policy optimization methods of Chapter 18;
  9. derive the direct preference optimization objective from a KL-regularized model;
  10. recognize reward hacking, length bias, and uncertainty as statistical modeling issues;
  11. implement toy reward-model and KL-regularized policy updates in Python;
  12. use AI tools to audit reward-model assumptions, preference data, and policy-update equations.

23.2 20.1 Why large models change the surface, not the mathematics

Modern language models are very large, but the main mathematical objects are familiar from earlier chapters. A model receives an input prompt \(x\) and produces a response

\[ y=(y_1,y_2,\ldots,y_T), \]

where each \(y_t\) is a token. A conditional language model defines a probability distribution

\[ \pi_\theta(y\mid x) = \prod_{t=1}^T \pi_\theta(y_t\mid x,y_1,\ldots,y_{t-1}). \]

This is a policy. The state at step \(t\) can be viewed as the partial transcript

\[ s_t=(x,y_1,\ldots,y_{t-1}), \]

and the action is the next token \(a_t=y_t\). The episode ends when an end-of-sequence token is produced or a maximum length is reached.

In practice, many alignment methods simplify the problem to a contextual bandit: the model chooses a whole response \(y\) for prompt \(x\) and receives a scalar score \(r(x,y)\). This suppresses token-level credit assignment and treats each response as one action.

For a mathematical first reading, use the contextual-bandit model. For algorithmic details such as PPO and token-level credit assignment, use the sequential MDP model.

The contextual-bandit abstraction is

\[ x\sim \rho,\qquad y\sim \pi_\theta(\cdot\mid x),\qquad R=r(x,y). \]

The objective is

\[ J(\theta)=\mathbb E_{x\sim \rho,\; y\sim \pi_\theta(\cdot\mid x)}[r(x,y)]. \]

This looks simpler than a full MDP, but the action space is enormous because \(y\) ranges over token sequences. Therefore the central issue is not writing down the objective; the central issue is estimating and optimizing it reliably.

The diagram emphasizes the statistical pipeline: collect preference data, fit a reward or preference model, optimize a policy with a regularization constraint, and audit the resulting policy.

23.3 20.2 Preference data

Human feedback is often collected as comparisons. For a prompt \(x\), suppose a human annotator sees two responses \(y^+\) and \(y^-\) and says that \(y^+\) is preferred to \(y^-\). The data have the form

\[ \mathcal D = \{(x_i,y_i^+,y_i^-)\}_{i=1}^n. \]

The superscript \(+\) means preferred and \(-\) means rejected. The preference label is not an absolute truth; it is a noisy observation. Two annotators may disagree, and the same annotator may be inconsistent across time. Thus preference learning is a statistical modeling problem.

A reward model assigns a scalar score

\[ r_\phi(x,y)\in \mathbb R. \]

The intended interpretation is that larger reward means stronger preference. However, only differences in reward are identifiable from pairwise comparisons. If we add a prompt-dependent constant \(c(x)\) to every response score, then

\[ r_\phi'(x,y)=r_\phi(x,y)+c(x) \]

produces the same preference probabilities under any model depending only on differences \(r_\phi(x,y^+)-r_\phi(x,y^-)\).

Identifiability up to prompt-dependent constants. Pairwise preferences for a fixed prompt identify differences between response scores, not absolute reward levels. Therefore \(r(x,y)\) and \(r(x,y)+c(x)\) are equivalent for pairwise preference modeling.

Reason. Pairwise models depend on \(r(x,y^+)-r(x,y^-)\). Adding \(c(x)\) to both terms cancels.

This matters because the reward model is later used for optimization. A reward model may have good pairwise accuracy but still contain biases or extrapolation errors that become dangerous under policy optimization.

23.4 20.3 Bradley-Terry preference model

A standard pairwise preference model is the Bradley-Terry model. It assumes

\[ \mathbb P(y^+\succ y^-\mid x) = \sigma\left(r_\phi(x,y^+)-r_\phi(x,y^-)\right), \]

where

\[ \sigma(z)=\frac{1}{1+\exp(-z)}. \]

Equivalently,

\[ \mathbb P(y^+\succ y^-\mid x) = \frac{\exp(r_\phi(x,y^+))} {\exp(r_\phi(x,y^+))+\exp(r_\phi(x,y^-))}. \]

Given data \(\mathcal D\), the negative log-likelihood is

\[ \mathcal L(\phi) = -\sum_{i=1}^n \log \sigma\left(r_\phi(x_i,y_i^+)-r_\phi(x_i,y_i^-)\right). \]

This is logistic regression on score differences. Let

\[ \Delta_i(\phi)=r_\phi(x_i,y_i^+)-r_\phi(x_i,y_i^-). \]

Then each observation contributes

\[ \ell_i(\phi)=-\log \sigma(\Delta_i(\phi)). \]

If the model is linear in features,

\[ r_\phi(x,y)=\phi^T f(x,y), \]

then

\[ \Delta_i(\phi) = \phi^T\left(f(x_i,y_i^+)-f(x_i,y_i^-)\right). \]

So preference learning becomes ordinary binary logistic regression on difference features.

The probability curve shows a key statistical point: very large reward differences saturate the preference probability. Once the model is confident, further increasing the reward difference changes the likelihood only slightly, but it may still strongly affect later policy optimization.

23.4.1 Python example: fitting a linear reward model

The following example creates synthetic pairwise preference data, fits a linear Bradley-Terry reward model, and evaluates whether the estimated reward direction agrees with the true reward direction.

import numpy as np

rng = np.random.default_rng(7243)

n = 1500
p = 4
true_phi = np.array([1.2, -0.7, 0.4, 0.9])

# Difference features d_i = f(x_i, y_i^+) - f(x_i, y_i^-).
# The sign is chosen so that positive true_phi^T d_i means the first response is preferred.
D_raw = rng.normal(size=(n, p))
logits = D_raw @ true_phi
prob_prefer_first = 1 / (1 + np.exp(-logits))
labels = rng.binomial(1, prob_prefer_first, size=n)

# Convert observations so D always means preferred minus rejected.
D = D_raw.copy()
D[labels == 0] *= -1

phi = np.zeros(p)
step = 0.15
for k in range(800):
    margins = D @ phi
    probs = 1 / (1 + np.exp(-margins))
    grad = -np.mean((1 - probs)[:, None] * D, axis=0)
    phi -= step * grad

cosine = phi @ true_phi / (np.linalg.norm(phi) * np.linalg.norm(true_phi))
print("estimated phi:", np.round(phi, 3))
print("true phi:     ", true_phi)
print("cosine agreement:", round(float(cosine), 3))
estimated phi: [ 1.278 -0.726  0.446  0.931]
true phi:      [ 1.2 -0.7  0.4  0.9]
cosine agreement: 1.0

The estimated vector does not need to equal the true vector exactly. In many preference models, only the direction and the induced ranking matter. In large-model alignment, the reward model is much richer than a linear model, but the same likelihood logic remains.

23.5 20.4 KL-regularized policy optimization

Suppose we already have a reward model \(r(x,y)\). A naive update would maximize

\[ \mathbb E_{x\sim \rho,\; y\sim \pi(\cdot\mid x)}[r(x,y)]. \]

This is risky because the optimizer may exploit reward-model errors. A standard remedy is to keep the new policy close to a reference policy \(\pi_0\), often the supervised fine-tuned model. The regularized objective is

\[ J_\beta(\pi) = \mathbb E_{x\sim \rho,\; y\sim \pi(\cdot\mid x)}[r(x,y)] - \beta\,\mathbb E_{x\sim \rho} \left[D_{\mathrm{KL}}(\pi(\cdot\mid x)\|\pi_0(\cdot\mid x))\right]. \]

For a fixed prompt \(x\), the optimization problem is

\[ \max_{p\in \Delta(Y)} \left\{ \sum_{y}p(y)r(x,y) - \beta\sum_y p(y)\log\frac{p(y)}{\pi_0(y\mid x)} \right\}. \]

The solution has a closed form.

KL-regularized optimizer. Assume \(Y\) is finite and \(\pi_0(y\mid x)>0\) for all \(y\). The maximizer of

\[ \sum_y p(y)r(x,y)-\beta\sum_y p(y)\log\frac{p(y)}{\pi_0(y\mid x)} \]

over \(p\in \Delta(Y)\) is

\[ p^*(y\mid x) = \frac{\pi_0(y\mid x)\exp(r(x,y)/\beta)} {Z_\beta(x)}, \]

where

\[ Z_\beta(x)=\sum_{z}\pi_0(z\mid x)\exp(r(x,z)/\beta). \]

The optimal value is

\[ \beta\log Z_\beta(x). \]

Proof. Introduce a Lagrange multiplier \(\lambda\) for \(\sum_y p(y)=1\). Differentiating with respect to \(p(y)\) gives

\[ r(x,y)-\beta\left(\log\frac{p(y)}{\pi_0(y\mid x)}+1\right)+\lambda=0. \]

Solving gives \(p(y)=C\pi_0(y\mid x)\exp(r(x,y)/\beta)\), and normalization determines \(C=1/Z_\beta(x)\).

This formula is one of the cleanest mathematical statements in RLHF. It says that the updated policy is the reference policy tilted by exponentiated reward. The parameter \(\beta\) controls the strength of the tilt.

When \(\beta\) is small, the policy concentrates on high-reward responses. When \(\beta\) is large, the policy remains closer to \(\pi_0\). Thus \(\beta\) plays the same conceptual role as an inverse trust-region strength.

23.5.1 Python example: exact KL-regularized response distribution

import numpy as np

responses = np.array(["short correct", "long helpful", "verbose unsafe", "refuses", "creative"])
pi0 = np.array([0.35, 0.25, 0.08, 0.22, 0.10])
reward = np.array([0.8, 1.4, 2.2, 0.1, 1.0])

def kl_regularized_policy(pi0, reward, beta):
    logits = np.log(pi0) + reward / beta
    logits -= logits.max()
    p = np.exp(logits)
    return p / p.sum()

for beta in [0.25, 0.75, 2.0]:
    p = kl_regularized_policy(pi0, reward, beta)
    print("beta =", beta)
    for name, prob in zip(responses, p):
        print(f"  {name:15s}: {prob:.3f}")
beta = 0.25
  short correct  : 0.014
  long helpful   : 0.110
  verbose unsafe : 0.866
  refuses        : 0.001
  creative       : 0.009
beta = 0.75
  short correct  : 0.213
  long helpful   : 0.339
  verbose unsafe : 0.315
  refuses        : 0.053
  creative       : 0.080
beta = 2.0
  short correct  : 0.314
  long helpful   : 0.303
  verbose unsafe : 0.145
  refuses        : 0.139
  creative       : 0.099

This small example also shows a danger. If the reward model gives a large score to a problematic response, then small \(\beta\) can amplify that mistake.

23.6 20.5 RLHF as PPO with a learned reward

A common RLHF training pipeline uses a reward model and then optimizes the language-model policy with PPO. At a high level:

  1. begin with a reference model \(\pi_0\);
  2. collect preference comparisons;
  3. fit a reward model \(r_\phi(x,y)\);
  4. sample responses from the current policy \(\pi_\theta\);
  5. compute a reward with a KL penalty;
  6. update \(\pi_\theta\) using PPO-style policy optimization.

For a sampled response \(y=(y_1,\ldots,y_T)\), the token-level log probability is

\[ \log \pi_\theta(y\mid x) = \sum_{t=1}^T \log \pi_\theta(y_t\mid x,y_{<t}). \]

A KL-penalized sequence-level reward can be written as

\[ \widetilde r(x,y) = r_\phi(x,y) - \beta\left(\log \pi_\theta(y\mid x)-\log \pi_0(y\mid x)\right). \]

The expression inside parentheses is a sampled estimate of a KL contribution. PPO then uses probability ratios

\[ r_t(\theta) = \frac{\pi_\theta(a_t\mid s_t)}{\pi_{\theta_{\mathrm{old}}}(a_t\mid s_t)} \]

and an advantage estimate \(\widehat A_t\) to build the clipped surrogate objective

\[ L^{\mathrm{CLIP}}(\theta) = \mathbb E_t \left[ \min\left( r_t(\theta)\widehat A_t, \operatorname{clip}(r_t(\theta),1-\epsilon,1+\epsilon)\widehat A_t \right) \right]. \]

This is mathematically the same policy optimization idea studied in Chapter 18. The difference is scale: the policy is a large neural sequence model and the reward is learned from preferences.

RLHF is not a new mathematical category. It combines preference-model estimation, KL-regularized control, and approximate policy-gradient optimization.

23.7 20.6 Direct preference optimization

Reward-model-based RLHF has two stages: estimate \(r_\phi\), then optimize a policy using that reward. Direct preference optimization, or DPO, uses the KL-regularized solution formula to train the policy directly from pairwise preference data.

From the KL-regularized optimizer,

\[ \pi^*(y\mid x) = \frac{1}{Z_\beta(x)}\pi_0(y\mid x)\exp(r(x,y)/\beta). \]

Rearranging gives

\[ r(x,y) = \beta\left( \log \pi^*(y\mid x)-\log \pi_0(y\mid x) \right) + \beta\log Z_\beta(x). \]

For two responses \(y^+\) and \(y^-\) to the same prompt, the normalizing term cancels:

\[ r(x,y^+)-r(x,y^-) = \beta \left[ \log\frac{\pi^*(y^+\mid x)}{\pi_0(y^+\mid x)} - \log\frac{\pi^*(y^-\mid x)}{\pi_0(y^-\mid x)} \right]. \]

DPO replaces the unknown optimal policy \(\pi^*\) by the parameterized policy \(\pi_\theta\). The DPO loss is

\[ \mathcal L_{\mathrm{DPO}}(\theta) = - \sum_{i=1}^n \log \sigma\left( \beta \left[ \log\frac{\pi_\theta(y_i^+\mid x_i)}{\pi_0(y_i^+\mid x_i)} - \log\frac{\pi_\theta(y_i^-\mid x_i)}{\pi_0(y_i^-\mid x_i)} \right] \right). \]

The expression inside the sigmoid is a difference of log-ratios. It rewards the policy for making preferred responses more likely than rejected responses relative to the reference model.

DPO can be read as logistic regression where the feature is the policy log-ratio difference. The reference policy still appears explicitly, so the method remains tied to KL-regularized optimization.

23.7.1 Python example: DPO loss for toy log probabilities

import numpy as np

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

def dpo_loss(logp_pos, logp_neg, logp0_pos, logp0_neg, beta=0.5):
    margin = beta * ((logp_pos - logp0_pos) - (logp_neg - logp0_neg))
    return -np.log(sigmoid(margin)), margin

examples = [
    (-2.1, -2.4, -2.2, -2.2),
    (-1.8, -2.7, -2.0, -2.1),
    (-2.5, -2.0, -2.1, -2.3),
]

for ex in examples:
    loss, margin = dpo_loss(*ex, beta=0.7)
    print("margin:", round(float(margin), 3), "loss:", round(float(loss), 3))
margin: 0.21 loss: 0.594
margin: 0.56 loss: 0.452
margin: -0.49 loss: 0.968

A positive margin means the policy prefers \(y^+\) over \(y^-\) more strongly than the reference model does. A negative margin means the policy is moving in the wrong direction for that comparison.

23.8 20.7 Reward hacking and over-optimization

A reward model is an estimator. If we optimize too aggressively against an imperfect estimator, the policy may discover responses that receive high reward but do not actually satisfy the intended human preference. This is called reward hacking or specification gaming.

A simple statistical model explains the issue. Suppose the learned reward is

\[ \widehat r(y)=r_{\mathrm{true}}(y)+\varepsilon(y), \]

where \(\varepsilon(y)\) is estimation error. If we choose

\[ y^*\in \arg\max_y \widehat r(y), \]

then we are likely to select a response with a large positive error term. This is an optimization version of selection bias.

Selection amplifies noise. If many candidate responses have noisy reward estimates, then maximizing the estimated reward tends to select candidates with positive noise. Therefore the expected true reward of the selected candidate can be much lower than its estimated reward.

KL regularization reduces this effect by limiting the search region around the reference policy. It does not eliminate reward hacking, but it reduces the chance that optimization pushes into poorly estimated regions.

The gap between selected estimated reward and selected true reward grows when the number of candidates increases or the reward noise grows.

23.9 20.8 Length bias and normalization

Large-model rewards can accidentally depend on superficial features such as response length. For example, if longer answers often receive higher preference labels in the training data, a reward model may learn

\[ \widehat r(x,y) \approx r_{\mathrm{content}}(x,y)+c\,\operatorname{length}(y). \]

If \(c>0\), optimization may produce unnecessarily long answers. If \(c<0\), the model may become terse or evasive. This is not merely a user-interface problem; it is a statistical confounding problem.

A useful diagnostic is to regress reward-model scores on response length and other simple covariates. A large coefficient does not prove bias, but it signals that the reward model may be using length as a proxy for quality.

For mathematical modeling, the lesson is that the reward model should be treated like any statistical estimator: inspect residuals, check covariates, estimate uncertainty, and test out-of-distribution behavior.

23.9.1 Python example: reward over-optimization under noisy estimates

import numpy as np

rng = np.random.default_rng(2026)

m = 2000
true_reward = rng.normal(loc=0.0, scale=1.0, size=m)
noise = rng.normal(loc=0.0, scale=0.8, size=m)
est_reward = true_reward + noise

chosen = np.argmax(est_reward)
print("chosen estimated reward:", round(float(est_reward[chosen]), 3))
print("chosen true reward:     ", round(float(true_reward[chosen]), 3))
print("average true reward:    ", round(float(true_reward.mean()), 3))
print("noise at chosen response:", round(float(noise[chosen]), 3))
chosen estimated reward: 4.772
chosen true reward:      3.334
average true reward:     -0.034
noise at chosen response: 1.439

In repeated simulations, the chosen response often has a large positive noise component. This is why strong optimization pressure can expose reward-model weaknesses.

23.10 20.9 Contextual-bandit gradient for sequence models

If the response \(y\) is treated as one action, then the score-function identity gives

\[ \nabla_\theta J(\theta) = \mathbb E_{x,y\sim \pi_\theta} \left[ \nabla_\theta\log \pi_\theta(y\mid x)\, r(x,y) \right]. \]

Because

\[ \log\pi_\theta(y\mid x) = \sum_{t=1}^T \log\pi_\theta(y_t\mid x,y_{<t}), \]

we have

\[ \nabla_\theta\log\pi_\theta(y\mid x) = \sum_{t=1}^T \nabla_\theta\log\pi_\theta(y_t\mid x,y_{<t}). \]

Thus a sequence-level reward assigns the same global signal to many token-level decisions. This creates a credit-assignment problem: which tokens were responsible for the final preference?

A baseline \(b(x)\) can reduce variance:

\[ \nabla_\theta J(\theta) = \mathbb E \left[ \nabla_\theta\log \pi_\theta(y\mid x)\, (r(x,y)-b(x)) \right]. \]

The reason this is valid is

\[ \mathbb E_{y\sim\pi_\theta(\cdot\mid x)} [\nabla_\theta\log \pi_\theta(y\mid x)b(x)] = b(x)\nabla_\theta\sum_y \pi_\theta(y\mid x) =0. \]

This is the same baseline identity used in Chapter 15.

23.11 20.10 Statistical issues in preference-based RL

Preference-based RL with large models involves several statistical assumptions.

23.11.1 Nonrandom data collection

Preference data are not usually sampled uniformly from all possible prompts and responses. They depend on prompt sources, model versions, sampling temperatures, annotator instructions, and filtering rules. The empirical distribution may differ from deployment.

Mathematically, the training risk is

\[ \mathbb E_{(x,y^+,y^-)\sim P_{\mathrm{train}}} [\ell_\phi(x,y^+,y^-)], \]

but deployment depends on a different distribution \(P_{\mathrm{deploy}}\). Generalization requires assumptions connecting these distributions.

23.11.2 Annotator noise and ambiguity

A pairwise label may not reflect a deterministic ordering. For some prompts, multiple responses may be valid. A probabilistic model such as

\[ \mathbb P(y_1\succ y_2\mid x)=\sigma(r(x,y_1)-r(x,y_2)) \]

should be interpreted as a model of noisy preferences, not a proof that a single scalar reward fully captures human judgment.

23.11.3 Reward uncertainty

A reward model can be confident in regions with many comparisons and unreliable in regions rarely seen during training. A conservative policy update should account for uncertainty. One conceptual form is

\[ \max_\pi \mathbb E_\pi[\widehat r(x,y)-\lambda u(x,y)] - \beta D_{\mathrm{KL}}(\pi\|\pi_0), \]

where \(u(x,y)\) is an uncertainty penalty.

For MS Statistics students, the reward model should be treated as an estimated statistical model. For MA Applied Math students, the policy update should be treated as a regularized optimization problem over probability measures.

23.12 20.11 AI-assisted learning components

AI-assisted derivation audit. Ask an AI system to derive the closed-form optimizer of the KL-regularized objective. Then check whether it correctly handles the normalization constant \(Z_\beta(x)\) and the sign of the KL penalty.

Preference-model critique. Give an AI system a small set of hypothetical preference pairs. Ask it to identify possible confounders such as length, tone, formatting, refusal style, or verbosity. Then translate those concerns into statistical covariates.

DPO code review. Provide an implementation of the DPO loss and ask an AI system to check whether the preferred and rejected log probabilities are in the correct order. Verify the answer manually by testing a case where the preferred response becomes more likely.

Reward hacking thought experiment. Ask an AI system to propose ways a model might exploit a reward model. Classify each failure as an estimation error, distribution-shift error, objective-misspecification error, or optimization error.

23.13 20.12 Summary

Reinforcement learning with large models is built from familiar mathematical parts:

  • a language model is a stochastic policy over token sequences;
  • preference data define a statistical estimation problem;
  • the Bradley-Terry model turns pairwise comparisons into logistic regression on reward differences;
  • KL-regularized policy optimization keeps the updated model close to a reference model;
  • the exact KL-regularized optimizer is a reward-tilted reference distribution;
  • PPO-based RLHF applies approximate policy-gradient optimization to a learned reward;
  • DPO uses the KL-regularized solution formula to train directly from preferences;
  • reward hacking, length bias, and distribution shift are statistical failures of the reward-learning-and-optimization pipeline.

The main lesson is that alignment-style RL is not only a deep-learning technique. It is a combination of probability, statistical estimation, convex duality, regularized optimization, and stochastic policy learning.

23.14 Exercises

23.14.1 Conceptual exercises

  1. Explain why a language model can be viewed as a policy. What are the states and actions in the sequential MDP view?
  2. Explain the difference between the contextual-bandit view and the token-level MDP view of response generation.
  3. Why are pairwise preferences insufficient to identify absolute reward levels?
  4. Why can a reward model with high pairwise accuracy still lead to poor policy optimization?
  5. Explain why KL regularization is useful when updating a large model.
  6. Compare reward-model-based RLHF and DPO. What mathematical object is eliminated in DPO?

23.14.2 Mathematical exercises

  1. Derive the Bradley-Terry probability formula from the logistic model \[ \mathbb P(y^+\succ y^-\mid x)=\sigma(r(x,y^+)-r(x,y^-)). \]
  2. Prove that adding \(c(x)\) to all rewards for a fixed prompt does not change pairwise preference probabilities.
  3. Derive the KL-regularized optimizer \[ p^*(y\mid x)=\frac{\pi_0(y\mid x)\exp(r(x,y)/\beta)}{Z_\beta(x)}. \]
  4. Show that the optimal value of the KL-regularized problem is \(\beta\log Z_\beta(x)\).
  5. Starting from the KL-regularized optimizer, derive the DPO log-ratio expression.
  6. Prove that subtracting a prompt-dependent baseline \(b(x)\) does not change the contextual-bandit policy gradient.

23.14.3 Computational exercises

  1. Simulate pairwise preference data from a known linear reward model and estimate the reward weights using gradient descent.
  2. For a fixed reference distribution \(\pi_0\) and reward vector \(r\), plot the KL-regularized optimizer for several values of \(\beta\).
  3. Implement the DPO loss for a batch of preferred and rejected log probabilities.
  4. Simulate reward hacking by adding noise to true rewards and selecting the response with highest estimated reward.
  5. Build a diagnostic that regresses reward-model scores on response length and flags a potential length bias.
  6. Compare a naive reward-maximizing policy with a KL-regularized policy in a finite response set.

23.14.4 AI-assisted exercises

  1. Ask an AI system to explain RLHF in mathematical language. Identify any missing assumptions.
  2. Ask an AI system to derive the DPO loss. Check whether the preferred and rejected responses appear in the correct order.
  3. Ask an AI system to list possible reward-model failure modes. Organize them as statistical errors, optimization errors, or deployment errors.
  4. Ask an AI system to write code for a Bradley-Terry reward model. Test it on synthetic data where the true answer is known.
  5. Ask an AI system to critique a reward function that gives higher scores to longer answers. Translate the critique into a measurable diagnostic.

23.15 Notes for instructors

For MA Applied Math students, emphasize the KL-regularized optimization problem, the Lagrange multiplier derivation, convex duality, and the connection with entropy-regularized control. For MS Statistics students, emphasize pairwise comparison models, identifiability, sampling bias, uncertainty, reward-model diagnostics, and distribution shift. The chapter can be taught as a bridge between reinforcement learning, statistical ranking, and modern AI alignment.

Recommended sequence for one lecture:

  1. language model as policy;
  2. pairwise preference data and Bradley-Terry likelihood;
  3. KL-regularized policy optimization and its closed-form solution;
  4. PPO-based RLHF and DPO as two algorithmic realizations;
  5. reward hacking and diagnostics.

23.16 References

The mathematical foundation of this chapter connects policy-gradient and PPO methods from Chapters 15 and 18 with preference-based alignment methods. Useful references include (schulman2017proximal?), (ouyang2022training?), (christiano2017deep?), (stiennon2020learning?), (ziegler2019fine?), and (rafailov2023direct?).