Core idea. Actor-critic methods combine two approximations. The actor is a parameterized policy \(\pi_\theta(a\mid s)\) that chooses actions. The critic is a value-function approximation, such as \(V_w(s)\) or \(Q_w(s,a)\), that evaluates the actor. The critic supplies a low-variance estimate of the advantage of the sampled action, and the actor uses that estimate to update the policy.
The REINFORCE algorithm replaces \(Q^{\pi_\theta}(S_t,A_t)\) by an observed return \(G_t\). This gives an unbiased estimator in many episodic settings, but the estimator may have high variance because a full trajectory return includes many random rewards unrelated to the sampled action.
Actor-critic methods reduce this variance by learning a critic. Instead of using the full return \(G_t\), the actor uses a value-based estimate of how much better the sampled action was than the policy’s typical behavior in that state. This estimate is often based on the temporal-difference error.
For MA Applied Math students, actor-critic is a coupled stochastic approximation method. For MS Statistics students, actor-critic is a sequential estimation method in which a nuisance function, the value function, is learned to reduce the variance of a policy-gradient estimator.
The critic does not replace the actor. It supplies information about the actor’s current policy. The actor changes the policy; the critic evaluates the policy being changed.
19.2.1 Interactive: Actor-critic information flow
The actor generates actions, the environment generates rewards and next states, and the critic turns the transition into a TD error. That scalar TD error updates both the critic and the actor.
19.3 16.2 The advantage function
For a fixed policy \(\pi\), the state-value and action-value functions are
19.8 16.7 Python example: exact advantages in a tiny MDP
We begin with a two-state episodic MDP. State \(2\) is terminal. In state \(0\), action \(1\) moves toward state \(1\), while action \(0\) tends to keep the process in state \(0\). In state \(1\), action \(1\) tends to reach the terminal reward.
A tabular actor-critic algorithm learns to move from state 0 to state 1 and then to the terminal reward.
Estimated values: [0.9237 0.9879 0. ]
Policy at state 0: [0.0711 0.9289]
Policy at state 1: [0.031 0.969]
The policy should learn to prefer action \(1\) in both nonterminal states. In state \(0\), action \(1\) moves the agent toward state \(1\). In state \(1\), action \(1\) moves the agent toward the terminal reward.
plt.figure(figsize=(7, 4))plt.plot(prob_good_0, label="Pr(action 1 | state 0)")plt.plot(prob_good_1, label="Pr(action 1 | state 1)")plt.xlabel("episode")plt.ylabel("probability")plt.title("Policy probabilities during actor-critic learning")plt.legend()plt.show()
The actor gradually increases the probability of useful actions.
19.10 16.9 Actor-critic as coupled stochastic approximation
The actor and critic are updated from the same trajectory, so actor-critic algorithms are coupled stochastic approximation schemes. A useful abstract form is
Here \(M_{t+1}^{(w)}\) and \(M_{t+1}^{(\theta)}\) are noise terms. Often the critic is trained on a faster time scale than the actor:
\[
\frac{\alpha_t}{\beta_t}\to 0.
\]
Intuitively, the critic nearly equilibrates for the current actor before the actor changes substantially. This separation is mathematically useful because the actor can then be analyzed as if it were receiving approximately correct value estimates.
19.10.1 Interactive: Two-time-scale learning
The critic usually needs to track the current policy faster than the actor changes it. When the actor moves too quickly, the critic may evaluate a policy that is already outdated.
19.11 16.10 \(n\)-step actor-critic
The one-step TD error uses a short bootstrap target. More generally, an \(n\)-step target is
Small \(n\) gives more bootstrapping and often lower variance. Large \(n\) uses more observed rewards and often lower bias when the critic is inaccurate.
19.12 16.11 Advantage actor-critic and batch updates
In modern implementations, one often collects a batch of transitions or short rollouts before updating the parameters. Suppose a batch contains samples indexed by \(i\). Let \(\widehat A_i\) be an advantage estimate. The actor objective is often written as
A batch update averages several noisy advantage estimates before changing the actor. Larger batches reduce gradient noise but require more data before each update.
19.13 16.12 Generalized advantage estimation
Generalized advantage estimation, or GAE, combines TD errors across multiple horizons. Define the one-step TD residual
The coefficient \(\eta>0\) controls how strongly the method values exploration. Too little entropy may cause early policy collapse. Too much entropy may prevent the policy from exploiting learned structure.
19.14.1 Interactive: Entropy-regularized actor update
Entropy regularization keeps action probabilities away from zero while the advantage term pushes probability toward useful actions.
19.15 16.14 Python example: GAE from a sequence of TD errors
The following function computes GAE for a finite rollout.
import numpy as npdef generalized_advantage_estimate(rewards, values, gamma=0.99, lam=0.95):"""Compute GAE for one finite rollout. rewards has length T. values has length T + 1 and includes the bootstrap value. """ T =len(rewards) advantages = np.zeros(T) gae =0.0for t inreversed(range(T)): delta = rewards[t] + gamma * values[t +1] - values[t] gae = delta + gamma * lam * gae advantages[t] = gaereturn advantagesrewards = np.array([0.0, 0.0, 1.0, 0.0, 2.0])values = np.array([0.30, 0.35, 0.60, 0.40, 1.10, 0.00])for lam in [0.0, 0.5, 0.95, 1.0]: adv = generalized_advantage_estimate(rewards, values, gamma=0.9, lam=lam)print(f"lambda={lam:0.2f}: {adv}")
With nonlinear function approximation, such as neural networks, the same equations are implemented by automatic differentiation. The mathematical meaning remains the same: the critic estimates a value baseline, and the actor increases the probability of actions with positive estimated advantage.
19.17 16.16 Stability issues
Actor-critic methods are powerful but delicate. Common sources of instability include:
the critic is inaccurate and gives misleading advantage estimates;
the actor changes too quickly for the critic to track;
the policy becomes nearly deterministic too early;
off-policy data are used without correction;
bootstrapping and function approximation amplify errors;
advantage estimates have high variance;
learning rates are poorly scaled.
Useful diagnostics include:
average episodic return;
critic loss or TD-error magnitude;
policy entropy;
gradient norms;
approximate KL divergence between old and new policies;
distribution of estimated advantages;
fraction of actions with near-zero probability.
Actor-critic should not be interpreted as a guaranteed gradient method unless the critic, sampling distribution, and update schedule satisfy additional assumptions. In practice, actor-critic is a family of approximate stochastic optimization methods.
19.18 16.17 AI-assisted learning components
19.18.1 AI prompt: deriving the actor update
Ask an AI assistant:
Starting from the policy gradient theorem, derive the one-step actor-critic update using the TD error as an advantage estimate. Explicitly state where an approximation is introduced.
and then explain that \(A^{\pi_\theta}\) is replaced by the sample TD error computed from an approximate critic.
19.18.2 AI prompt: debugging actor-critic code
Ask an AI assistant to check the following implementation points:
Are terminal states handled correctly in the TD target?
Is the actor using \(\nabla_\theta\log\pi_\theta(A_t\mid S_t)\) rather than \(\nabla_\theta\pi_\theta(A_t\mid S_t)\)?
Is the advantage estimate detached from the actor computation when using neural-network libraries?
Are actor and critic learning rates separated?
Is entropy being maximized rather than minimized?
19.18.3 AI prompt: experiment critique
Give an AI assistant a plot of returns, value losses, entropy, and gradient norms. Ask it to propose three possible failure modes and three follow-up experiments. The answer should distinguish evidence from speculation.
19.19 16.18 Summary
Actor-critic methods combine policy-gradient optimization with value-function estimation. The actor changes the policy, while the critic estimates the quality of the current policy. The central bridge between them is the TD error:
When the critic is accurate, the conditional expectation of \(\delta_t\) is the advantage \(A^\pi(s,a)\). This makes the TD error a natural low-variance substitute for a Monte Carlo return in the policy-gradient update.
The main mathematical themes are:
score-function gradients;
baselines and advantage functions;
Bellman equations;
temporal-difference learning;
stochastic approximation;
two-time-scale dynamics;
bias-variance tradeoffs in advantage estimation.
19.20 Exercises
19.20.1 Conceptual exercises
Explain the difference between an actor and a critic.
Why is \(V^\pi(s)\) a natural baseline for policy gradients?
Explain why the TD error is not exactly the advantage when the critic is approximate.
Why can a critic reduce variance but introduce bias?
Explain why entropy regularization can help exploration.
Modify the tabular actor-critic example by changing \(\alpha\) and \(\beta\). What happens when \(\alpha\) is much larger than \(\beta\)?
Add entropy regularization to the tabular actor update.
Replace the one-step TD error by a three-step advantage estimate.
Plot the distribution of TD errors during learning.
Compare actor-critic with REINFORCE on the same tiny MDP.
19.20.4 AI-assisted exercises
Ask an AI assistant to derive the actor-critic update. Then identify one step in the derivation that relies on approximation.
Ask an AI assistant to inspect actor-critic code and identify terminal-state bugs.
Ask an AI assistant to design diagnostics for unstable actor-critic training.
Ask an AI assistant to explain the difference between A2C, A3C, and PPO in mathematical terms.
Ask an AI assistant to critique whether a proposed reward function is likely to produce unintended behavior.
19.21 Notes for instructors
This chapter is a natural synthesis point. Students have already seen policy gradients, TD learning, and stochastic approximation. The lecture can be organized around one identity:
when the critic is exact. This identity makes actor-critic methods feel inevitable rather than mysterious.
For MA Applied Math students, emphasize coupled stochastic approximation, fixed points, and two-time-scale dynamics. For MS Statistics students, emphasize baselines, variance reduction, conditional expectations, nuisance estimation, and diagnostics.