Core idea. Entropy-regularized reinforcement learning replaces the hard goal of always choosing the currently best action by a smoother optimization problem that rewards both performance and controlled randomness. The central mathematical identity is the convex duality formula
where \(H(p)=-\sum_a p(a)\log p(a)\). This identity transforms Bellman optimality equations, policy gradients, and actor-critic methods into their soft versions.
22.1 Learning goals
After reading this chapter, students should be able to:
define entropy for a discrete policy and interpret it as a measure of policy randomness;
explain why exploration is a mathematical difficulty rather than only a programming detail;
derive the log-sum-exp formula as the entropy-regularized maximum;
compute the softmax optimizer associated with an entropy-regularized action choice;
write the soft Bellman optimality equations for finite discounted MDPs;
prove that the soft Bellman optimality operator is a contraction;
distinguish ordinary optimal control from maximum-entropy control;
explain soft policy iteration and soft Q-learning;
describe the objective behind soft actor-critic;
implement small soft-RL examples in Python;
use AI tools to audit entropy terms, temperature parameters, and soft Bellman equations.
22.2 19.1 Why exploration needs mathematics
In previous chapters, we studied control methods such as SARSA, Q-learning, and policy-gradient learning. These methods must decide not only what to exploit but also what to explore. The basic tension is this:
choosing actions that currently look good may produce high reward now;
choosing uncertain actions may reveal better behavior later.
For a finite MDP, suppose an agent has an estimate \(Q(s,a)\) of action values. A purely greedy policy chooses
This policy can stop exploring too early. If the estimate \(Q(s,a)\) is inaccurate, greedy behavior may repeatedly choose a suboptimal action. This is especially important in sample-based learning, where data are generated by the policy itself.
where \(a^*(s)\) is a greedy action. This is useful, but it is not derived from an optimization principle. Entropy-regularized RL gives a cleaner mathematical answer: choose a policy that maximizes expected value plus a reward for randomness.
Entropy-regularized RL changes the optimization problem. Instead of asking only for the action with largest estimated value, it asks for a probability distribution over actions that balances large value and large entropy.
22.3 19.2 Entropy of a policy
Let \(A\) be a finite action set and let \(p\) be a probability distribution on \(A\). The Shannon entropy of \(p\) is
\[
H(p)=-\sum_{a\in A}p(a)\log p(a),
\]
with the convention \(0\log 0=0\).
For a policy \(\pi\), the entropy at state \(s\) is
Entropy is small when the policy is nearly deterministic and large when the policy is spread over many actions. If \(|A|=m\), then
\[
0\le H(\pi(\cdot\mid s))\le \log m.
\]
The maximum is achieved by the uniform distribution \(\pi(a\mid s)=1/m\).
Entropy maximum on a finite action set. Let \(A\) have \(m\) actions. Among all probability vectors \(p\in\Delta(A)\), entropy is maximized by the uniform distribution \(p(a)=1/m\). The maximum value is \(\log m\).
Proof sketch. Since the function \(x\mapsto -x\log x\) is concave, entropy is concave on the probability simplex. Using Lagrange multipliers for the constraint \(\sum_a p(a)=1\) gives \(-\log p(a)-1=\lambda\), so all positive coordinates are equal. Hence \(p(a)=1/m\) and \(H(p)=\log m\).
For two actions, write \(p\) for the probability of action \(1\). Then
\[
H(p)=-p\log p-(1-p)\log(1-p).
\]
This function is symmetric around \(p=1/2\), equals zero at \(p=0\) and \(p=1\), and reaches its maximum \(\log 2\) at \(p=1/2\).
The graph shows that entropy rewards uncertainty. A deterministic policy has no entropy bonus, while a balanced randomized policy has the largest bonus.
22.4 19.3 Entropy-regularized objectives
In a discounted finite MDP, the ordinary policy objective is
The term \(-\log\pi(A_t\mid S_t)\) is large when the policy assigns small probability to the sampled action. Entropy regularization therefore discourages the policy from collapsing too quickly to a single action.
The coefficient \(\alpha\) changes the problem being solved. With \(\alpha=0\), the objective is ordinary expected return. With large \(\alpha\), the agent may prefer high-entropy behavior even when it sacrifices reward.
22.5 19.4 The softmax distribution from optimization
Consider a single state with action values \(q(a)\). Instead of choosing an action by the hard maximum
\[
\max_{a\in A} q(a),
\]
we choose a probability distribution \(p\in\Delta(A)\) by solving
The second term is \(-\alpha D_{\mathrm{KL}}(p\|p^*)\le 0\). Equality holds exactly when \(p=p^*\).
This theorem explains why softmax policies are natural in reinforcement learning. They are not merely a convenient numerical trick; they solve an entropy-regularized optimization problem.
As \(\alpha\downarrow 0\), the softmax distribution concentrates on maximizers of \(q\). As \(\alpha\uparrow\infty\), the distribution approaches uniform.
is called the temperature-scaled log-sum-exp. It is a smooth approximation to the maximum. If \(m=|A|\), then
\[
\max_a q(a)
\le
\operatorname{LSE}_\alpha(q)
\le
\max_a q(a)+\alpha\log m.
\]
The lower bound follows because the sum of exponentials is at least the largest exponential. The upper bound follows because the sum of exponentials is at most \(m\) times the largest exponential.
This inequality gives a precise interpretation of the temperature. The approximation error between the soft maximum and the hard maximum is at most \(\alpha\log m\).
The derivative of log-sum-exp with respect to \(q(a)\) is the softmax probability:
In ordinary Bellman optimality, the action selection step uses \(\max_a\). In soft Bellman optimality, this is replaced by \(\alpha\log\sum_a\exp(\cdot/\alpha)\).
The improvement step is no longer greedy in the hard sense. It is greedy with respect to the entropy-regularized objective. This produces a policy that favors high-value actions but retains controlled randomness.
Soft policy iteration is especially useful as a conceptual bridge:
Ordinary RL
Soft RL
hard maximum
log-sum-exp soft maximum
deterministic greedy policy
softmax policy
Bellman optimality
soft Bellman optimality
maximize expected reward
maximize reward plus entropy
exploration added externally
exploration built into objective
22.10 19.9 Python example: entropy and softmax
The following code computes entropy, softmax probabilities, and log-sum-exp values for a fixed vector of action values.
import numpy as npq = np.array([-1.0, 0.0, 1.0, 2.0])def softmax(q, alpha=1.0): z = (q - np.max(q)) / alpha e = np.exp(z)return e / e.sum()def entropy(p): p = np.asarray(p, dtype=float) positive = p >0return-np.sum(p[positive] * np.log(p[positive]))def logsumexp_alpha(q, alpha=1.0): m = np.max(q)return m + alpha * np.log(np.sum(np.exp((q - m) / alpha)))for alpha in [0.1, 0.3, 1.0, 3.0]: p = softmax(q, alpha)print(f"alpha = {alpha:3.1f}")print("policy:", np.round(p, 3))print("entropy:", round(entropy(p), 3))print("soft value:", round(logsumexp_alpha(q, alpha), 3))print()
The policy is stochastic even after convergence. It assigns larger probability to actions with larger soft action values, but it does not collapse completely unless \(\alpha\) is very small.
22.12 19.11 The limit \(\alpha\to 0\)
Soft RL contains ordinary optimal control as a limiting case. Since
the soft Bellman equation converges to the ordinary Bellman optimality equation. Also, the softmax policy converges to a distribution supported on the maximizing actions.
When the maximizer is unique, the limiting policy is deterministic. When several actions tie, the limiting policy may distribute probability over the tied actions depending on the limiting procedure.
Soft actor-critic, or SAC, is a modern off-policy actor-critic method based on the maximum-entropy objective. In continuous-control problems, the objective is often written as
This objective encourages the actor to assign high probability to actions with large Q-value, but penalizes policies that become too concentrated.
The temperature \(\alpha\) controls the strength of entropy. A common adaptive-temperature objective is designed to push the policy entropy toward a target value. Conceptually:
if entropy is too low, increase \(\alpha\) to encourage exploration;
if entropy is too high, decrease \(\alpha\) to focus more on reward.
22.15 19.14 Statistical viewpoint
Entropy regularization has a natural statistical interpretation. It changes a deterministic optimization problem into a smoother estimation problem.
First, softmax probabilities make action selection less sensitive to small estimation errors in \(Q(s,a)\). If two actions have almost equal estimated values, a greedy policy may switch discontinuously when the estimates change slightly. A softmax policy changes smoothly.
Second, entropy can improve data collection. If the policy remains stochastic, the agent observes more state-action pairs. This can reduce estimation bias caused by narrow data coverage.
Third, entropy introduces a bias-variance tradeoff. A larger \(\alpha\) usually increases exploration and may reduce estimation variance, but it also changes the target policy away from the reward-maximizing deterministic policy.
For MS Statistics students, entropy regularization should be viewed as a smoothing and data-collection mechanism. It makes the policy less brittle, improves coverage, and modifies the statistical target.
22.16 19.15 AI-assisted learning components
AI derivation prompt: entropy-regularized maximum
Ask an AI assistant:
Derive the identity \(\max_{p\in\Delta(A)} \sum_a p(a)q(a)+\alpha H(p)=\alpha\log\sum_a\exp(q(a)/\alpha)\). Use KL divergence in the proof and explain why the maximizer is the softmax distribution.
Then check whether the assistant correctly handles the constraint \(\sum_a p(a)=1\) and the temperature \(\alpha\).
AI code-review prompt: stable log-sum-exp
Ask an AI assistant:
Review this function for numerical stability when \(q/\alpha\) is large. Explain why subtracting the maximum before exponentiating is necessary.
Use the assistant’s answer to rewrite a naive implementation of log-sum-exp.
AI modeling prompt: choosing \(\alpha\)
Ask an AI assistant:
In an entropy-regularized reinforcement learning problem, what symptoms indicate that \(\alpha\) is too large or too small? Give diagnostics using return, entropy, action probabilities, and state-action coverage.
Then classify the diagnostics as mathematical, statistical, or computational.
AI comparison prompt: PPO entropy bonus versus SAC
Ask an AI assistant:
Compare the entropy bonus often added to PPO with the maximum-entropy objective used in SAC. Which one is usually an auxiliary regularizer, and which one is central to the Bellman equations?
Check whether the answer distinguishes on-policy and off-policy learning.
22.17 19.16 Summary
Entropy-regularized RL replaces hard maximization by soft maximization. The central identity is
The soft Bellman operator remains a contraction, so the fixed-point logic from dynamic programming still applies. Entropy regularization also provides a principled way to connect exploration, convex duality, stochastic policies, and modern algorithms such as soft actor-critic.
22.18 Exercises
22.18.1 Conceptual exercises
Explain why a greedy policy can fail when the value estimates are inaccurate.
Explain the difference between adding random exploration externally and optimizing an entropy-regularized objective.
Describe what happens to a softmax policy as \(\alpha\downarrow 0\) and as \(\alpha\uparrow\infty\).
Explain why entropy regularization changes the objective being optimized.
Compare \(\epsilon\)-greedy exploration and softmax exploration.
Explain why log-sum-exp is a smooth approximation to the maximum.
22.18.2 Mathematical exercises
Prove that entropy is maximized by the uniform distribution on a finite set.
For a two-action policy, compute the derivative and second derivative of \(H(p)=-p\log p-(1-p)\log(1-p)\).
Prove the entropy-regularized maximization theorem using Lagrange multipliers.
Derive the soft Bellman optimality equation from the entropy-regularized one-step optimization problem.
Prove that the soft Bellman optimality operator is a \(\gamma\)-contraction.
Show that the gradient of log-sum-exp is the softmax distribution.
Derive the fixed-policy soft Bellman equation for \(V_\alpha^\pi\).
22.18.3 Computational exercises
Implement entropy and softmax for a vector of action values. Plot the action probabilities as \(\alpha\) varies.
Implement soft value iteration for a finite MDP and compare it with ordinary value iteration.
For a fixed MDP, plot \(\|V_\alpha^*-V^*\|_\infty\) as a function of \(\alpha\).
Implement soft Q-learning on a small gridworld.
Compare the state-action visitation frequencies under greedy, \(\epsilon\)-greedy, and softmax policies.
Implement stable and unstable versions of log-sum-exp and identify when the unstable version fails.
Simulate adaptive temperature updates that increase \(\alpha\) when entropy is below a target and decrease it when entropy is above a target.
22.18.4 AI-assisted exercises
Ask an AI assistant to explain the difference between entropy regularization and exploration noise. Critique the response.
Ask an AI assistant to derive the soft Bellman equation. Check whether it uses \(-\alpha\log\pi(a\mid s)\) with the correct sign.
Give an AI assistant a naive log-sum-exp implementation and ask it to find numerical stability issues.
Ask an AI assistant to compare DQN, PPO, and SAC from the viewpoint of the Bellman operator used by each method.
Ask an AI assistant to design diagnostics for deciding whether a learned SAC policy has too much or too little entropy.
22.19 Notes for instructors
This chapter is a natural transition from policy optimization to modern maximum-entropy RL. For MA Applied Math students, emphasize convex duality, log-sum-exp smoothing, contraction mappings, and soft Bellman fixed points. For MS Statistics students, emphasize entropy as smoothing, policy stochasticity as data coverage, and temperature as a bias-variance-control parameter.
A possible lecture sequence is:
entropy and two-action examples;
entropy-regularized maximization and the softmax theorem;