22  Entropy, Exploration, and Soft RL

Core idea. Entropy-regularized reinforcement learning replaces the hard goal of always choosing the currently best action by a smoother optimization problem that rewards both performance and controlled randomness. The central mathematical identity is the convex duality formula

\[ \max_{p\in\Delta(A)}\left\{\sum_{a\in A}p(a)q(a)+\alpha H(p)\right\} = \alpha\log\sum_{a\in A}\exp\left(\frac{q(a)}{\alpha}\right), \]

where \(H(p)=-\sum_a p(a)\log p(a)\). This identity transforms Bellman optimality equations, policy gradients, and actor-critic methods into their soft versions.

22.1 Learning goals

After reading this chapter, students should be able to:

  1. define entropy for a discrete policy and interpret it as a measure of policy randomness;
  2. explain why exploration is a mathematical difficulty rather than only a programming detail;
  3. derive the log-sum-exp formula as the entropy-regularized maximum;
  4. compute the softmax optimizer associated with an entropy-regularized action choice;
  5. write the soft Bellman optimality equations for finite discounted MDPs;
  6. prove that the soft Bellman optimality operator is a contraction;
  7. distinguish ordinary optimal control from maximum-entropy control;
  8. explain soft policy iteration and soft Q-learning;
  9. describe the objective behind soft actor-critic;
  10. implement small soft-RL examples in Python;
  11. use AI tools to audit entropy terms, temperature parameters, and soft Bellman equations.

22.2 19.1 Why exploration needs mathematics

In previous chapters, we studied control methods such as SARSA, Q-learning, and policy-gradient learning. These methods must decide not only what to exploit but also what to explore. The basic tension is this:

  • choosing actions that currently look good may produce high reward now;
  • choosing uncertain actions may reveal better behavior later.

For a finite MDP, suppose an agent has an estimate \(Q(s,a)\) of action values. A purely greedy policy chooses

\[ \pi_{\text{greedy}}(a\mid s) = \mathbf 1\left\{a\in \arg\max_b Q(s,b)\right\}. \]

This policy can stop exploring too early. If the estimate \(Q(s,a)\) is inaccurate, greedy behavior may repeatedly choose a suboptimal action. This is especially important in sample-based learning, where data are generated by the policy itself.

A simple exploration rule is \(\epsilon\)-greedy:

\[ \pi_\epsilon(a\mid s) = (1-\epsilon)\mathbf 1\{a=a^*(s)\}+\frac{\epsilon}{|A|}, \]

where \(a^*(s)\) is a greedy action. This is useful, but it is not derived from an optimization principle. Entropy-regularized RL gives a cleaner mathematical answer: choose a policy that maximizes expected value plus a reward for randomness.

Entropy-regularized RL changes the optimization problem. Instead of asking only for the action with largest estimated value, it asks for a probability distribution over actions that balances large value and large entropy.

22.3 19.2 Entropy of a policy

Let \(A\) be a finite action set and let \(p\) be a probability distribution on \(A\). The Shannon entropy of \(p\) is

\[ H(p)=-\sum_{a\in A}p(a)\log p(a), \]

with the convention \(0\log 0=0\).

For a policy \(\pi\), the entropy at state \(s\) is

\[ H(\pi(\cdot\mid s)) = -\sum_{a\in A}\pi(a\mid s)\log \pi(a\mid s). \]

Entropy is small when the policy is nearly deterministic and large when the policy is spread over many actions. If \(|A|=m\), then

\[ 0\le H(\pi(\cdot\mid s))\le \log m. \]

The maximum is achieved by the uniform distribution \(\pi(a\mid s)=1/m\).

Entropy maximum on a finite action set. Let \(A\) have \(m\) actions. Among all probability vectors \(p\in\Delta(A)\), entropy is maximized by the uniform distribution \(p(a)=1/m\). The maximum value is \(\log m\).

Proof sketch. Since the function \(x\mapsto -x\log x\) is concave, entropy is concave on the probability simplex. Using Lagrange multipliers for the constraint \(\sum_a p(a)=1\) gives \(-\log p(a)-1=\lambda\), so all positive coordinates are equal. Hence \(p(a)=1/m\) and \(H(p)=\log m\).

For two actions, write \(p\) for the probability of action \(1\). Then

\[ H(p)=-p\log p-(1-p)\log(1-p). \]

This function is symmetric around \(p=1/2\), equals zero at \(p=0\) and \(p=1\), and reaches its maximum \(\log 2\) at \(p=1/2\).

The graph shows that entropy rewards uncertainty. A deterministic policy has no entropy bonus, while a balanced randomized policy has the largest bonus.

22.4 19.3 Entropy-regularized objectives

In a discounted finite MDP, the ordinary policy objective is

\[ J(\pi) = \mathbb E_\pi\left[\sum_{t=0}^{\infty}\gamma^t R_{t+1}\right]. \]

The entropy-regularized objective adds a policy entropy term:

\[ J_\alpha(\pi) = \mathbb E_\pi\left[ \sum_{t=0}^{\infty}\gamma^t \left( R_{t+1} +\alpha H(\pi(\cdot\mid S_t)) \right) \right], \]

where \(\alpha>0\) is the temperature or entropy coefficient.

Equivalently, because

\[ H(\pi(\cdot\mid s)) = \mathbb E_{A\sim\pi(\cdot\mid s)}[-\log\pi(A\mid s)], \]

the entropy-regularized reward can be written along a trajectory as

\[ R_{t+1}-\alpha\log \pi(A_t\mid S_t). \]

Thus the objective is

\[ J_\alpha(\pi) = \mathbb E_\pi\left[ \sum_{t=0}^{\infty}\gamma^t \left( R_{t+1}-\alpha\log \pi(A_t\mid S_t) \right) \right]. \]

The term \(-\log\pi(A_t\mid S_t)\) is large when the policy assigns small probability to the sampled action. Entropy regularization therefore discourages the policy from collapsing too quickly to a single action.

The coefficient \(\alpha\) changes the problem being solved. With \(\alpha=0\), the objective is ordinary expected return. With large \(\alpha\), the agent may prefer high-entropy behavior even when it sacrifices reward.

22.5 19.4 The softmax distribution from optimization

Consider a single state with action values \(q(a)\). Instead of choosing an action by the hard maximum

\[ \max_{a\in A} q(a), \]

we choose a probability distribution \(p\in\Delta(A)\) by solving

\[ \max_{p\in\Delta(A)} \left\{ \sum_a p(a)q(a)+\alpha H(p) \right\}. \]

This problem has a closed-form solution.

Entropy-regularized maximization. For \(\alpha>0\) and a finite action set \(A\),

\[ \max_{p\in\Delta(A)} \left\{ \sum_a p(a)q(a)+\alpha H(p) \right\} = \alpha\log\sum_a \exp\left(\frac{q(a)}{\alpha}\right). \]

The unique maximizer is

\[ p^*(a) = \frac{\exp(q(a)/\alpha)}{\sum_b \exp(q(b)/\alpha)}. \]

Proof. Define

\[ Z=\sum_b \exp(q(b)/\alpha) \]

and

\[ p^*(a)=\frac{\exp(q(a)/\alpha)}{Z}. \]

Then

\[ q(a)=\alpha\log p^*(a)+\alpha\log Z. \]

For any distribution \(p\),

\[ \sum_a p(a)q(a)+\alpha H(p) = \alpha\log Z - \alpha\sum_a p(a)\log\frac{p(a)}{p^*(a)}. \]

The second term is \(-\alpha D_{\mathrm{KL}}(p\|p^*)\le 0\). Equality holds exactly when \(p=p^*\).

This theorem explains why softmax policies are natural in reinforcement learning. They are not merely a convenient numerical trick; they solve an entropy-regularized optimization problem.

As \(\alpha\downarrow 0\), the softmax distribution concentrates on maximizers of \(q\). As \(\alpha\uparrow\infty\), the distribution approaches uniform.

22.6 19.5 Log-sum-exp as a smooth maximum

The value

\[ \operatorname{LSE}_\alpha(q) = \alpha\log\sum_a\exp(q(a)/\alpha) \]

is called the temperature-scaled log-sum-exp. It is a smooth approximation to the maximum. If \(m=|A|\), then

\[ \max_a q(a) \le \operatorname{LSE}_\alpha(q) \le \max_a q(a)+\alpha\log m. \]

The lower bound follows because the sum of exponentials is at least the largest exponential. The upper bound follows because the sum of exponentials is at most \(m\) times the largest exponential.

This inequality gives a precise interpretation of the temperature. The approximation error between the soft maximum and the hard maximum is at most \(\alpha\log m\).

The derivative of log-sum-exp with respect to \(q(a)\) is the softmax probability:

\[ \frac{\partial}{\partial q(a)} \operatorname{LSE}_\alpha(q) = \frac{\exp(q(a)/\alpha)}{\sum_b\exp(q(b)/\alpha)}. \]

Thus the soft value is smooth, and its gradient is a probability distribution over actions.

22.7 19.6 Soft value functions

For a fixed policy \(\pi\), define the entropy-regularized value function

\[ V_\alpha^\pi(s) = \mathbb E_\pi\left[ \sum_{t=0}^{\infty}\gamma^t \left( R_{t+1}-\alpha\log\pi(A_t\mid S_t) \right) \mid S_0=s \right]. \]

The corresponding action-value function is

\[ Q_\alpha^\pi(s,a) = \mathbb E_\pi\left[ \sum_{t=0}^{\infty}\gamma^t \left( R_{t+1}-\alpha\log\pi(A_t\mid S_t) \right) \mid S_0=s,A_0=a \right]. \]

For finite MDPs, these functions satisfy

\[ Q_\alpha^\pi(s,a) = r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V_\alpha^\pi(s'), \]

and

\[ V_\alpha^\pi(s) = \sum_a\pi(a\mid s) \left[ Q_\alpha^\pi(s,a)-\alpha\log\pi(a\mid s) \right]. \]

Equivalently, the fixed-policy soft Bellman operator is

\[ (T_\alpha^\pi V)(s) = \sum_a\pi(a\mid s) \left[ r(s,a) -\alpha\log\pi(a\mid s) +\gamma\sum_{s'}P(s'\mid s,a)V(s') \right]. \]

This is still a contraction in the sup norm:

\[ \|T_\alpha^\pi V-T_\alpha^\pi W\|_\infty \le \gamma\|V-W\|_\infty. \]

The entropy term changes the reward, but it does not change the discount contraction factor.

22.8 19.7 The soft Bellman optimality equations

For optimal entropy-regularized control, define

\[ V_\alpha^*(s)=\sup_\pi V_\alpha^\pi(s). \]

The soft optimal action-value function satisfies

\[ Q_\alpha^*(s,a) = r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V_\alpha^*(s'). \]

The optimal soft value is obtained by entropy-regularized maximization over actions:

\[ V_\alpha^*(s) = \alpha\log\sum_a\exp\left(\frac{Q_\alpha^*(s,a)}{\alpha}\right). \]

The optimal soft policy is

\[ \pi_\alpha^*(a\mid s) = \frac{\exp(Q_\alpha^*(s,a)/\alpha)}{\sum_b\exp(Q_\alpha^*(s,b)/\alpha)}. \]

In ordinary Bellman optimality, the action selection step uses \(\max_a\). In soft Bellman optimality, this is replaced by \(\alpha\log\sum_a\exp(\cdot/\alpha)\).

Define the soft Bellman optimality operator

\[ (T_\alpha V)(s) = \alpha\log\sum_a\exp\left( \frac{r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V(s')}{\alpha} \right). \]

Soft Bellman contraction. For \(\alpha>0\), the operator \(T_\alpha\) is a \(\gamma\)-contraction under the sup norm:

\[ \|T_\alpha V-T_\alpha W\|_\infty \le \gamma\|V-W\|_\infty. \]

Therefore, \(T_\alpha\) has a unique fixed point \(V_\alpha^*\), and soft value iteration \(V_{k+1}=T_\alpha V_k\) converges to \(V_\alpha^*\).

Proof sketch. For each state \(s\), define

\[ x_a=r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V(s') \]

and

\[ y_a=r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)W(s'). \]

Then \(|x_a-y_a|\le\gamma\|V-W\|_\infty\) for all actions. The log-sum-exp function is 1-Lipschitz with respect to the sup norm. Hence

\[ |(T_\alpha V)(s)-(T_\alpha W)(s)| \le \gamma\|V-W\|_\infty. \]

Taking the maximum over states gives the result.

22.9 19.8 Soft policy iteration

Soft policy iteration alternates between two steps.

Soft policy iteration

  1. Soft policy evaluation. Given \(\pi_k\), solve or approximate

    \[ V_\alpha^{\pi_k}=T_\alpha^{\pi_k}V_\alpha^{\pi_k}. \]

  2. Soft policy improvement. Compute

    \[ Q_\alpha^{\pi_k}(s,a) = r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V_\alpha^{\pi_k}(s') \]

    and update

    \[ \pi_{k+1}(a\mid s) \propto \exp\left(\frac{Q_\alpha^{\pi_k}(s,a)}{\alpha}\right). \]

The improvement step is no longer greedy in the hard sense. It is greedy with respect to the entropy-regularized objective. This produces a policy that favors high-value actions but retains controlled randomness.

Soft policy iteration is especially useful as a conceptual bridge:

Ordinary RL Soft RL
hard maximum log-sum-exp soft maximum
deterministic greedy policy softmax policy
Bellman optimality soft Bellman optimality
maximize expected reward maximize reward plus entropy
exploration added externally exploration built into objective

22.10 19.9 Python example: entropy and softmax

The following code computes entropy, softmax probabilities, and log-sum-exp values for a fixed vector of action values.

import numpy as np

q = np.array([-1.0, 0.0, 1.0, 2.0])

def softmax(q, alpha=1.0):
    z = (q - np.max(q)) / alpha
    e = np.exp(z)
    return e / e.sum()

def entropy(p):
    p = np.asarray(p, dtype=float)
    positive = p > 0
    return -np.sum(p[positive] * np.log(p[positive]))

def logsumexp_alpha(q, alpha=1.0):
    m = np.max(q)
    return m + alpha * np.log(np.sum(np.exp((q - m) / alpha)))

for alpha in [0.1, 0.3, 1.0, 3.0]:
    p = softmax(q, alpha)
    print(f"alpha = {alpha:3.1f}")
    print("policy:", np.round(p, 3))
    print("entropy:", round(entropy(p), 3))
    print("soft value:", round(logsumexp_alpha(q, alpha), 3))
    print()
alpha = 0.1
policy: [0. 0. 0. 1.]
entropy: 0.0
soft value: 2.0

alpha = 0.3
policy: [0.    0.001 0.034 0.964]
entropy: 0.16
soft value: 2.011

alpha = 1.0
policy: [0.032 0.087 0.237 0.644]
entropy: 0.948
soft value: 2.44

alpha = 3.0
policy: [0.142 0.198 0.276 0.385]
entropy: 1.32
soft value: 4.864

For small \(\alpha\), the policy is close to greedy. For large \(\alpha\), the policy is more diffuse and entropy is larger.

22.11 19.10 Python example: soft value iteration

We now implement soft value iteration for a tiny finite MDP with three states and two actions.

import numpy as np

num_states = 3
num_actions = 2
gamma = 0.90
alpha = 0.50

# P[s, a, sp]
P = np.array([
    [[0.75, 0.25, 0.00], [0.10, 0.80, 0.10]],
    [[0.20, 0.60, 0.20], [0.00, 0.25, 0.75]],
    [[0.10, 0.20, 0.70], [0.00, 0.10, 0.90]],
])

R = np.array([
    [0.0, 0.4],
    [0.2, 0.8],
    [0.1, 1.0],
])

def stable_logsumexp(x):
    m = np.max(x)
    return m + np.log(np.sum(np.exp(x - m)))

def soft_bellman(V, alpha):
    Q = R + gamma * np.einsum("sak,k->sa", P, V)
    return np.array([alpha * stable_logsumexp(Q[s] / alpha) for s in range(num_states)])

def soft_policy_from_Q(Q, alpha):
    logits = Q / alpha
    logits = logits - logits.max(axis=1, keepdims=True)
    e = np.exp(logits)
    return e / e.sum(axis=1, keepdims=True)

V = np.zeros(num_states)
history = []
for k in range(80):
    V_new = soft_bellman(V, alpha)
    history.append(np.max(np.abs(V_new - V)))
    V = V_new

Q = R + gamma * np.einsum("sak,k->sa", P, V)
pi = soft_policy_from_Q(Q, alpha)

print("Soft optimal value:", np.round(V, 3))
print("Soft optimal Q:")
print(np.round(Q, 3))
print("Soft optimal policy:")
print(np.round(pi, 3))
print("Final Bellman update size:", history[-1])
Soft optimal value: [ 9.711 10.269 10.469]
Soft optimal Q:
[[ 8.866  9.61 ]
 [ 9.378 10.177]
 [ 9.418 10.404]]
Soft optimal policy:
[[0.184 0.816]
 [0.168 0.832]
 [0.122 0.878]]
Final Bellman update size: 0.0002530186696496628

The policy is stochastic even after convergence. It assigns larger probability to actions with larger soft action values, but it does not collapse completely unless \(\alpha\) is very small.

22.12 19.11 The limit \(\alpha\to 0\)

Soft RL contains ordinary optimal control as a limiting case. Since

\[ \lim_{\alpha\downarrow 0} \alpha\log\sum_a\exp(q(a)/\alpha) = \max_a q(a), \]

the soft Bellman equation converges to the ordinary Bellman optimality equation. Also, the softmax policy converges to a distribution supported on the maximizing actions.

When the maximizer is unique, the limiting policy is deterministic. When several actions tie, the limiting policy may distribute probability over the tied actions depending on the limiting procedure.

alphas = [2.0, 1.0, 0.5, 0.2, 0.05]
q_state = np.array([0.0, 1.0, 1.2])

for alpha in alphas:
    p = softmax(q_state, alpha)
    soft_value = logsumexp_alpha(q_state, alpha)
    print(f"alpha={alpha:4.2f}  p={np.round(p, 3)}  soft value={soft_value:.3f}")

print("hard max:", np.max(q_state))
alpha=2.00  p=[0.224 0.369 0.408]  soft value=2.995
alpha=1.00  p=[0.142 0.386 0.472]  soft value=1.951
alpha=0.50  p=[0.052 0.381 0.568]  soft value=1.483
alpha=0.20  p=[0.002 0.268 0.73 ]  soft value=1.263
alpha=0.05  p=[0.    0.018 0.982]  soft value=1.201
hard max: 1.2

This limit is useful mathematically, but in computation very small \(\alpha\) can cause numerical instability and premature loss of exploration.

22.13 19.12 Soft Q-learning

The ordinary Q-learning target is

\[ R_{t+1}+\gamma\max_{a'}Q(S_{t+1},a'). \]

Soft Q-learning replaces the hard maximum by log-sum-exp:

\[ R_{t+1} + \gamma\alpha\log\sum_{a'}\exp\left(\frac{Q(S_{t+1},a')}{\alpha}\right). \]

The update becomes

\[ Q(S_t,A_t) \leftarrow Q(S_t,A_t) + \eta_t \left[ R_{t+1} + \gamma V(S_{t+1}) - Q(S_t,A_t) \right], \]

where

\[ V(s)=\alpha\log\sum_a\exp(Q(s,a)/\alpha). \]

The induced policy is

\[ \pi(a\mid s) =\frac{\exp(Q(s,a)/\alpha)}{\sum_b\exp(Q(s,b)/\alpha)}. \]

This gives an off-policy value-based method whose learned policy is stochastic.

rng = np.random.default_rng(19)

num_states = 3
num_actions = 2
gamma = 0.90
alpha = 0.40
eta = 0.08

Q = np.zeros((num_states, num_actions))
state = 0

# Small simulator using P and R from the previous example.
def step(s, a):
    sp = rng.choice(num_states, p=P[s, a])
    reward = R[s, a] + rng.normal(0, 0.05)
    return sp, reward

def soft_value_from_Q(Q, s, alpha):
    return logsumexp_alpha(Q[s], alpha)

def behavior_action(Q, s, alpha):
    p = softmax(Q[s], alpha)
    return rng.choice(num_actions, p=p)

errors = []
for t in range(5000):
    action = behavior_action(Q, state, alpha)
    next_state, reward = step(state, action)
    target = reward + gamma * soft_value_from_Q(Q, next_state, alpha)
    Q[state, action] += eta * (target - Q[state, action])
    state = next_state
    if (t + 1) % 500 == 0:
        policy = np.vstack([softmax(Q[s], alpha) for s in range(num_states)])
        errors.append(np.max(policy, axis=1).mean())

print("Learned soft Q:")
print(np.round(Q, 3))
print("Induced soft policy:")
print(np.round(np.vstack([softmax(Q[s], alpha) for s in range(num_states)]), 3))
Learned soft Q:
[[0.036 0.093]
 [0.124 9.517]
 [0.393 9.782]]
Induced soft policy:
[[0.464 0.536]
 [0.    1.   ]
 [0.    1.   ]]

22.14 19.13 Maximum-entropy actor-critic and SAC

Soft actor-critic, or SAC, is a modern off-policy actor-critic method based on the maximum-entropy objective. In continuous-control problems, the objective is often written as

\[ J_\alpha(\pi) = \mathbb E_\pi\left[ \sum_{t=0}^{\infty}\gamma^t \left( R_{t+1}-\alpha\log\pi(A_t\mid S_t) \right) \right]. \]

SAC learns:

  1. one or more soft Q-functions \(Q_\phi(s,a)\);
  2. a stochastic actor \(\pi_\theta(a\mid s)\);
  3. sometimes a temperature \(\alpha\) that adapts to match a target entropy.

The soft Q target has the form

\[ y = r+ \gamma \mathbb E_{a'\sim\pi_\theta(\cdot\mid s')} \left[ Q_{\bar\phi}(s',a')- \alpha\log\pi_\theta(a'\mid s') \right]. \]

The actor update approximately solves

\[ \min_\theta \mathbb E_{s\sim D,\,a\sim\pi_\theta} \left[ \alpha\log\pi_\theta(a\mid s)-Q_\phi(s,a) \right]. \]

This objective encourages the actor to assign high probability to actions with large Q-value, but penalizes policies that become too concentrated.

The temperature \(\alpha\) controls the strength of entropy. A common adaptive-temperature objective is designed to push the policy entropy toward a target value. Conceptually:

  • if entropy is too low, increase \(\alpha\) to encourage exploration;
  • if entropy is too high, decrease \(\alpha\) to focus more on reward.

22.15 19.14 Statistical viewpoint

Entropy regularization has a natural statistical interpretation. It changes a deterministic optimization problem into a smoother estimation problem.

First, softmax probabilities make action selection less sensitive to small estimation errors in \(Q(s,a)\). If two actions have almost equal estimated values, a greedy policy may switch discontinuously when the estimates change slightly. A softmax policy changes smoothly.

Second, entropy can improve data collection. If the policy remains stochastic, the agent observes more state-action pairs. This can reduce estimation bias caused by narrow data coverage.

Third, entropy introduces a bias-variance tradeoff. A larger \(\alpha\) usually increases exploration and may reduce estimation variance, but it also changes the target policy away from the reward-maximizing deterministic policy.

For MS Statistics students, entropy regularization should be viewed as a smoothing and data-collection mechanism. It makes the policy less brittle, improves coverage, and modifies the statistical target.

22.16 19.15 AI-assisted learning components

AI derivation prompt: entropy-regularized maximum

Ask an AI assistant:

Derive the identity \(\max_{p\in\Delta(A)} \sum_a p(a)q(a)+\alpha H(p)=\alpha\log\sum_a\exp(q(a)/\alpha)\). Use KL divergence in the proof and explain why the maximizer is the softmax distribution.

Then check whether the assistant correctly handles the constraint \(\sum_a p(a)=1\) and the temperature \(\alpha\).

AI code-review prompt: stable log-sum-exp

Ask an AI assistant:

Review this function for numerical stability when \(q/\alpha\) is large. Explain why subtracting the maximum before exponentiating is necessary.

Use the assistant’s answer to rewrite a naive implementation of log-sum-exp.

AI modeling prompt: choosing \(\alpha\)

Ask an AI assistant:

In an entropy-regularized reinforcement learning problem, what symptoms indicate that \(\alpha\) is too large or too small? Give diagnostics using return, entropy, action probabilities, and state-action coverage.

Then classify the diagnostics as mathematical, statistical, or computational.

AI comparison prompt: PPO entropy bonus versus SAC

Ask an AI assistant:

Compare the entropy bonus often added to PPO with the maximum-entropy objective used in SAC. Which one is usually an auxiliary regularizer, and which one is central to the Bellman equations?

Check whether the answer distinguishes on-policy and off-policy learning.

22.17 19.16 Summary

Entropy-regularized RL replaces hard maximization by soft maximization. The central identity is

\[ \max_p \left\{\langle p,q\rangle+\alpha H(p)\right\} =\alpha\log\sum_a\exp(q(a)/\alpha), \]

with optimizer

\[ p^*(a)\propto \exp(q(a)/\alpha). \]

This gives the soft Bellman equations

\[ V_\alpha^*(s) =\alpha\log\sum_a\exp\left(\frac{Q_\alpha^*(s,a)}{\alpha}\right), \]

and

\[ Q_\alpha^*(s,a) =r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V_\alpha^*(s'). \]

The soft Bellman operator remains a contraction, so the fixed-point logic from dynamic programming still applies. Entropy regularization also provides a principled way to connect exploration, convex duality, stochastic policies, and modern algorithms such as soft actor-critic.

22.18 Exercises

22.18.1 Conceptual exercises

  1. Explain why a greedy policy can fail when the value estimates are inaccurate.
  2. Explain the difference between adding random exploration externally and optimizing an entropy-regularized objective.
  3. Describe what happens to a softmax policy as \(\alpha\downarrow 0\) and as \(\alpha\uparrow\infty\).
  4. Explain why entropy regularization changes the objective being optimized.
  5. Compare \(\epsilon\)-greedy exploration and softmax exploration.
  6. Explain why log-sum-exp is a smooth approximation to the maximum.

22.18.2 Mathematical exercises

  1. Prove that entropy is maximized by the uniform distribution on a finite set.

  2. For a two-action policy, compute the derivative and second derivative of \(H(p)=-p\log p-(1-p)\log(1-p)\).

  3. Prove the entropy-regularized maximization theorem using Lagrange multipliers.

  4. Prove the inequality

    \[ \max_a q(a) \le \alpha\log\sum_a\exp(q(a)/\alpha) \le \max_a q(a)+\alpha\log |A|. \]

  5. Derive the soft Bellman optimality equation from the entropy-regularized one-step optimization problem.

  6. Prove that the soft Bellman optimality operator is a \(\gamma\)-contraction.

  7. Show that the gradient of log-sum-exp is the softmax distribution.

  8. Derive the fixed-policy soft Bellman equation for \(V_\alpha^\pi\).

22.18.3 Computational exercises

  1. Implement entropy and softmax for a vector of action values. Plot the action probabilities as \(\alpha\) varies.
  2. Implement soft value iteration for a finite MDP and compare it with ordinary value iteration.
  3. For a fixed MDP, plot \(\|V_\alpha^*-V^*\|_\infty\) as a function of \(\alpha\).
  4. Implement soft Q-learning on a small gridworld.
  5. Compare the state-action visitation frequencies under greedy, \(\epsilon\)-greedy, and softmax policies.
  6. Implement stable and unstable versions of log-sum-exp and identify when the unstable version fails.
  7. Simulate adaptive temperature updates that increase \(\alpha\) when entropy is below a target and decrease it when entropy is above a target.

22.18.4 AI-assisted exercises

  1. Ask an AI assistant to explain the difference between entropy regularization and exploration noise. Critique the response.
  2. Ask an AI assistant to derive the soft Bellman equation. Check whether it uses \(-\alpha\log\pi(a\mid s)\) with the correct sign.
  3. Give an AI assistant a naive log-sum-exp implementation and ask it to find numerical stability issues.
  4. Ask an AI assistant to compare DQN, PPO, and SAC from the viewpoint of the Bellman operator used by each method.
  5. Ask an AI assistant to design diagnostics for deciding whether a learned SAC policy has too much or too little entropy.

22.19 Notes for instructors

This chapter is a natural transition from policy optimization to modern maximum-entropy RL. For MA Applied Math students, emphasize convex duality, log-sum-exp smoothing, contraction mappings, and soft Bellman fixed points. For MS Statistics students, emphasize entropy as smoothing, policy stochasticity as data coverage, and temperature as a bias-variance-control parameter.

A possible lecture sequence is:

  1. entropy and two-action examples;
  2. entropy-regularized maximization and the softmax theorem;
  3. log-sum-exp as a smooth maximum;
  4. soft Bellman equations and contraction;
  5. soft policy iteration and soft Q-learning;
  6. SAC as maximum-entropy actor-critic;
  7. diagnostics and AI-assisted code review.

Recommended references include (sutton2018reinforcement?) for exploration and RL foundations, (haarnoja2018soft?) for SAC, and (puterman1994markov?) for the dynamic-programming background.