Core idea. A deep Q-network, or DQN, replaces the tabular action-value table \(Q(s,a)\) by a neural approximation \(Q_ heta(s,a)\). The target is still the Bellman optimality equation,
but the fixed point is approximated by stochastic optimization over sampled transitions. The two stabilizing devices in the original DQN idea are experience replay and a target network. The mathematical viewpoint is
This works when the state and action spaces are small enough to store a table. Many modern applications have high-dimensional states: images, text embeddings, sensor vectors, market features, or large simulation states. In that case, a table is not practical.
DQN replaces the table by a parameterized function
\[
Q_ heta:S\times A\to \mathbb R,
\]
where \(\theta\) denotes all neural-network weights. For a finite action space, one common architecture maps a state \(s\) to a vector of action values:
The key mathematical change is that a single parameter vector \(\theta\) controls many state-action values simultaneously. One gradient update at one transition changes the predicted values of other states and actions as well.
DQN should not be viewed as a new Bellman equation. It is an approximate method for solving the same Bellman optimality equation, but in a nonlinear function class.
20.2.1 Interactive: From state to action values
This diagram shows the usual DQN architecture: a state vector is mapped through a feature representation to one scalar value for each discrete action.
20.3 17.2 The Bellman target
For a transition
\[
(s,a,r,s'),
\]
the tabular Q-learning target is
\[
y
=
r+
\gamma \max_{a'} Q(s',a').
\]
In DQN, the current prediction is \(Q_\theta(s,a)\). A first attempt would use
However, this uses the same parameters both to define the target and to fit the prediction. As \(\theta\) changes, the target moves. This can create instability.
DQN introduces a separate target parameter vector \(\theta^-\). The target becomes
\[
y
=
r+
\gamma\max_{a'}Q_{\theta^-}(s',a').
\]
The online network \(Q_\theta\) is fitted to this temporarily fixed target. The target network is updated more slowly.
For terminal next states, the future term is omitted:
\[
y
=
r
\quad \text{if } s' \text{ is terminal}.
\]
A convenient way to write both cases is
\[
y
=
r+\gamma(1-d)\max_{a'}Q_{\theta^-}(s',a'),
\]
where \(d=1\) for terminal transitions and \(d=0\) otherwise.
20.4 17.3 The DQN loss
Let \(D\) be a distribution over stored transitions. In practice, \(D\) is the empirical distribution induced by the replay buffer. The DQN objective is the expected squared temporal-difference error
This is the same algebraic structure as semi-gradient TD: the target is not differentiated through.
If one differentiates through the max target using the same network, the update is no longer the standard DQN semi-gradient update. The standard method treats the target as fixed for the current optimization step.
20.4.1 Interactive: Squared TD loss and Huber loss
Large TD errors can dominate a squared-error objective. Many DQN implementations use the Huber loss to reduce sensitivity to extreme errors.
20.5 17.4 Target networks
The target network is a slowly updated copy of the online network. In a hard update, one sets
\[
\theta^-\leftarrow \theta
\]
every \(C\) training steps. Between such updates, \(\theta^-\) is fixed.
where \(0<\tau\ll 1\). This is also called Polyak averaging.
The target network helps because it slows down the movement of the regression target. If the target changes too quickly, the optimization problem changes while the optimizer is trying to solve it. This creates a moving-target problem.
20.5.1 Interactive: Target-network lag
The online network moves every step. The target network either jumps occasionally or tracks smoothly by averaging.
20.6 17.5 Experience replay
DQN stores observed transitions in a replay buffer:
\[
B
=
\{(s_i,a_i,r_i,s_i',d_i)\}_{i=1}^N.
\]
At each training step, a mini-batch is sampled from this buffer. Replay has several mathematical roles.
First, it improves data efficiency. A transition can be used more than once.
Second, it weakens temporal correlation. Sequential RL data are dependent:
\[
S_t,A_t,R_{t+1},S_{t+1},A_{t+1},\ldots
\]
are not iid samples. Random mini-batches from a buffer are still not perfectly iid, but they are less correlated than consecutive transitions.
Third, replay changes the training distribution. The loss is not weighted by the current policy alone. It is weighted by the empirical distribution of the buffer, which is a mixture of past behavior policies.
Let \(\mu_B(s,a)\) denote the empirical state-action distribution in the buffer. Then DQN approximately minimizes a projected Bellman-error-like objective under the weighting \(\mu_B\):
The buffer distribution matters. If important state-action pairs are rarely stored, the network may fit them poorly.
20.6.1 Interactive: Replay buffer age distribution
A replay buffer stores a moving window of transitions. Uniform sampling gives different probabilities to recent and old transitions depending on the buffer capacity.
Take a gradient step on \[
\frac{1}{m}\sum_{i=1}^m\left(y_i-Q_\theta(s_i,a_i)\right)^2.
\]
Every \(C\) steps, update \(\theta^-\leftarrow\theta\).
The policy used to collect data is usually \(\epsilon\)-greedy:
\[
A_t
=
\begin{cases}
\text{a random action}, & \text{with probability } \epsilon_t,\\
\arg\max_a Q_\theta(S_t,a), & \text{with probability } 1-\epsilon_t.
\end{cases}
\]
Thus DQN is an off-policy algorithm: the target is greedy, while the behavior policy includes exploration.
20.8 17.7 Python example: replay buffer and DQN target
The following example implements a small replay buffer and computes DQN targets from a linear action-value model. The point is to make the data structure and target calculation transparent before using neural-network libraries.
up to a constant factor depending on whether the loss includes \(1/2\).
alpha =0.05W = W_online.copy()b = b_online.copy()q_before = q_values(states, W, b)[np.arange(len(actions)), actions]loss_before = np.mean((targets - q_before) **2)# Use the gradient of 1/2 times the squared TD error.for s, a, y inzip(states, actions, targets): pred = s @ W[:, a] + b[a] delta = y - pred W[:, a] += alpha * delta * s b[a] += alpha * deltaq_after = q_values(states, W, b)[np.arange(len(actions)), actions]loss_after = np.mean((targets - q_after) **2)print("loss before:", round(float(loss_before), 5))print("loss after: ", round(float(loss_after), 5))
loss before: 0.16655
loss after: 0.07794
This small example is not yet deep learning, but it contains the mathematical structure of a DQN update. A neural-network optimizer replaces the explicit linear update by backpropagation.
20.10 17.9 Neural networks as action-value approximators
A neural DQN can be written abstractly as
\[
Q_\theta(s,a)
=
\left[f_\theta(s)\right]_a,
\]
where \(f_\theta(s)\in\mathbb R^{|A|}\) is a vector-valued neural network.
and the target \(y_i\) is treated as fixed during the online-network update.
For image inputs, the feature map may be a convolutional network. For vector states, it may be a multilayer perceptron. For structured or text states, it may involve embeddings or a transformer encoder. The Bellman target remains the same; the representation changes.
In DQN, the neural network is not the source of the objective. The objective comes from the Bellman optimality equation. The neural network defines the approximation class used to represent action values.
20.11 17.10 The deadly triad
Deep Q-learning combines three ingredients:
function approximation: \(Q_\theta\) is not tabular;
bootstrapping: targets depend on current or target value estimates;
off-policy learning: the behavior policy differs from the greedy target policy.
Together, these form the deadly triad. The term refers to the fact that the combination can produce instability or divergence.
Mathematically, the tabular Bellman optimality operator is a contraction in the sup norm. Once we restrict to a nonlinear function class and train by stochastic gradients under a replay-buffer distribution, we are no longer simply applying a contraction operator to a value table. Instead, the algorithm alternates between approximate target construction and approximate regression:
\[
Q_{\theta_k}
\quad \longrightarrow \quad
\text{targets from } Q_{\theta_k^-}
\quad \longrightarrow \quad
\theta_{k+1} \text{ by regression}.
\]
The regression step is influenced by optimization error, approximation error, sampling error, and distribution shift.
A small training loss does not guarantee an optimal policy. The loss is measured on the replay-buffer distribution and against bootstrapped targets, not directly against \(Q^*\).
20.12 17.11 Overestimation bias and Double DQN
Q-learning uses a maximum over estimated action values:
\[
\max_a \widehat Q(s,a).
\]
If the estimates contain noise, then the maximum tends to be biased upward. For example, if
\[
\widehat Q(s,a)=Q(s,a)+\varepsilon_a,
\]
with zero-mean errors \(\varepsilon_a\), then often
This is useful computationally, but it also raises statistical questions. If exploration decays too quickly, some actions may never be sampled enough. If exploration remains too large, the learned behavior may be inefficient.
DQN can also use noisy networks, entropy-based exploration, bootstrapped ensembles, count-based bonuses, or uncertainty-based exploration. Those methods are beyond the scope of this introductory chapter, but they all address the same mathematical problem: learning good action values requires enough coverage of state-action space.
20.14 17.13 Training diagnostics
DQN training curves can be difficult to interpret. Common diagnostics include:
episode return;
average loss;
average absolute TD error;
max action value;
replay-buffer coverage;
policy entropy or fraction of random actions;
target-network update times;
gradient norm.
A decreasing TD loss is not sufficient evidence of success. The learned policy may still be poor if the replay buffer is narrow, the targets are biased, or the network generalizes incorrectly.
20.14.1 Interactive: DQN diagnostics
A stable-looking loss does not always imply improving return. Multiple diagnostics should be read together.
20.15 17.14 Python example: a tiny DQN-style training loop
The following toy example trains a linear action-value model from replayed transitions generated by a simple environment. It is not intended to be a high-performance RL implementation. It is designed to show the mathematical ingredients of DQN in a small setting.
The state is a two-dimensional vector. The agent chooses one of two actions. Action \(1\) tends to increase the first coordinate and earns reward when the first coordinate is positive. Action \(0\) tends to decrease it.
import numpy as nprng = np.random.default_rng(2026)state_dim =2n_actions =2capacity =500batch_size =32gamma =0.95alpha =0.03epsilon =0.20target_update =25n_steps =300buffer= ReplayBuffer(capacity=capacity, state_dim=state_dim)W = rng.normal(scale=0.1, size=(state_dim, n_actions))b = np.zeros(n_actions)W_targ = W.copy()b_targ = b.copy()s = rng.normal(scale=0.3, size=state_dim)returns = []losses = []episode_return =0.0def env_step(s, a, rng): drift = np.array([-0.12, 0.00]) if a ==0else np.array([0.12, 0.00]) sp =0.90* s + drift + rng.normal(scale=0.08, size=state_dim) r =1.0if sp[0] >0.45else0.0 done =bool(abs(sp[0]) >1.2or rng.random() <0.02)return sp, r, donefor t inrange(n_steps):if rng.random() < epsilon: a =int(rng.integers(n_actions))else: a =int(np.argmax(q_values(s.reshape(1, -1), W, b)[0])) sp, r, done = env_step(s, a, rng)buffer.add(s, a, r, sp, done) episode_return += rifbuffer.size >= batch_size: states, actions, rewards, next_states, done_flags =buffer.sample(batch_size, rng) q_next = q_values(next_states, W_targ, b_targ) y = rewards + gamma * (1.0- done_flags) * np.max(q_next, axis=1) q_pred = q_values(states, W, b)[np.arange(batch_size), actions] delta = y - q_pred losses.append(float(np.mean(delta **2)))for si, ai, di inzip(states, actions, delta): W[:, ai] += alpha * di * si b[ai] += alpha * diif (t +1) % target_update ==0: W_targ = W.copy() b_targ = b.copy()if done: returns.append(episode_return) episode_return =0.0 s = rng.normal(scale=0.3, size=state_dim)else: s = spprint("number of completed episodes:", len(returns))print("mean return over last 10 episodes:", round(float(np.mean(returns[-10:])), 3) if returns elseNone)print("mean recent loss:", round(float(np.mean(losses[-20:])), 5) if losses elseNone)print("learned weights:")print(np.round(W, 3))print("learned biases:", np.round(b, 3))
number of completed episodes: 10
mean return over last 10 episodes: 18.9
mean recent loss: 1.16994
learned weights:
[[ 2.785 2.798]
[-0.025 -0.077]]
learned biases: [3.704 4.233]
This example contains the key DQN pattern:
collect a transition using an exploratory behavior policy;
store it in replay;
sample a mini-batch from replay;
compute targets using a target network;
update only the online network;
periodically copy the online network to the target network.
20.16 17.15 DQN as approximate value iteration
Value iteration applies the optimal Bellman operator:
\[
Q_{k+1}
=
TQ_k.
\]
DQN can be informally viewed as approximate value iteration:
where \(\Pi_{\mu_B}\) denotes projection or regression under the replay-buffer distribution. This notation should be read cautiously because the function class is nonlinear and the optimization may not find a global projection. Still, it captures the structure:
optimization error from incomplete gradient descent;
approximation error from the neural function class;
distribution error from using replay-buffer weighting;
target error from bootstrapping with an imperfect target network.
This decomposition is useful when debugging a DQN system.
20.17 17.16 Practical implementation checklist
Checklist for a DQN implementation
Confirm that the action space is discrete.
Normalize or scale state features when appropriate.
Use a replay buffer large enough to diversify mini-batches.
Do not train before the buffer contains enough transitions.
Use a separate target network.
Handle terminal transitions correctly by removing the future term.
Track return, TD error, loss, gradient norm, and action frequencies.
Compare with a random policy and a simple heuristic baseline.
Test the target calculation on a tiny hand-checkable batch.
Set random seeds when running diagnostic experiments.
20.18 17.17 AI-assisted learning components
AI prompt: equation audit
Ask an AI assistant to check whether the following target is correct:
\[
y_i=r_i+\gamma\max_{a'}Q_\theta(s_i',a').
\]
Then ask it to explain why the standard DQN target usually uses \(Q_{\theta^-}\) instead of \(Q_\theta\), and why terminal transitions require a factor \(1-d_i\).
AI prompt: code review
Provide your replay-buffer sampling code and ask the AI assistant to check for the following mistakes: sampling uninitialized entries, using replacement unintentionally, mixing terminal and nonterminal targets incorrectly, and updating the target network before computing the target.
AI prompt: experiment critique
Give the assistant a plot of return, TD loss, average max-\(Q\), and epsilon. Ask it to list at least three possible explanations for a decreasing loss but flat return. Then ask which extra diagnostic it would collect next.
20.19 17.18 Summary
DQN is the first major deep RL method in this book. Its mathematical core is not mysterious: it is Q-learning with a neural action-value approximation. The central target is
\[
y
=
r+\gamma(1-d)\max_{a'}Q_{\theta^-}(s',a'),
\]
and the online network minimizes a regression loss toward this target. Experience replay and target networks are not cosmetic tricks; they are attempts to control statistical dependence and target instability. Double DQN addresses maximization bias by separating action selection and action evaluation.
For applied mathematics students, DQN is approximate dynamic programming with nonlinear approximation and moving targets. For statistics students, DQN is sequential supervised learning with dependent data, biased targets, distribution shift, and exploration-driven sampling.
20.20 Conceptual exercises
Explain why a DQN needs one output per action when the action space is finite and discrete.
Why does DQN use a target network? Describe the moving-target problem in your own words.
Explain how experience replay changes the dependence structure of the training data.
Why is DQN considered off-policy?
Give an example where a small TD loss might not imply a good policy.
Explain the deadly triad and identify which three parts appear in DQN.
Compare DQN with tabular Q-learning. What is gained and what is lost?
Why does a terminal transition require removing the bootstrap term?
20.21 Mathematical exercises
Derive the gradient of \[
\frac{1}{2}\left(y-Q_\theta(s,a)\right)^2
\] with respect to \(\theta\), treating \(y\) as constant.
Suppose \(Q_\theta(s,a)=s^Tw_a+b_a\). Derive the update for \(w_a\) and \(b_a\) under the half-squared TD loss.
Prove that if \(\widehat Q_a=Q_a+\varepsilon_a\) with zero-mean noise, then \[
\mathbb E\max_a \widehat Q_a
\ge
\max_a Q_a.
\] Hint: use Jensen’s inequality and the convexity of the maximum function.
Let the soft target update be \[
\theta^-_{t+1}=(1-\tau)\theta^-_t+\tau\theta_t.
\] If \(\theta_t=\theta\) is constant for all \(t\), solve for \(\theta^-_t\).
In a mini-batch with terminal indicators \(d_i\), show how the DQN targets can be written as a vectorized NumPy expression.
Explain why the expression \[
Q_{\theta^-}\left(s',\arg\max_{a'}Q_\theta(s',a')\right)
\] separates action selection from action evaluation.
20.22 Computational exercises
Modify the toy DQN example by changing the replay capacity. How do the recent losses and returns change?
Replace the hard target update with a soft target update. Compare training behavior for several values of \(\tau\).
Implement the Huber loss and compare it with the squared loss when occasional large rewards are added.
Implement Double DQN targets in the toy linear example.
Track action frequencies over training. Does the learned policy collapse to one action too early?
Build a small finite MDP and compare tabular Q-learning with a linear DQN using one-hot state features.
Create a diagnostic plot with three curves: return, TD loss, and average max-\(Q\).
20.23 AI-assisted exercises
Ask an AI assistant to generate a minimal DQN implementation. Then manually check whether it uses a target network and handles terminal states correctly.
Ask an AI assistant to explain the difference between DQN and Double DQN. Identify any vague or incorrect statements.
Provide your own training diagnostics to an AI assistant and ask for possible failure modes. Separate evidence-based comments from speculation.
Ask an AI assistant to convert a tabular Q-learning implementation to DQN. Check whether the resulting code still assumes a table anywhere.
Ask an AI assistant to design a unit test for the target calculation. Implement and run the test.
20.24 Notes for instructors
This chapter is a bridge from tabular RL to deep RL. It should not be taught as a collection of software tricks. The clean mathematical sequence is:
For MA Applied Math students, emphasize approximate fixed points, nonlinear approximation, and the loss of contraction guarantees. For MS Statistics students, emphasize dependent data, replay-buffer distributions, target bias, and diagnostic uncertainty.