Roadmap

A learning path through The Mathematics of Reinforcement Learning

Roadmap

Note

This roadmap explains how the chapters, labs, mathematical ideas, and projects fit together. It can be used for a semester course, an accelerated module, or independent study.

Repository

The GitHub repository for the book is:

https://github.com/wanghemath/Book-MathRL

The labs are stored in:

labs/

The final integrated project is:

labs/chapter-28-lab.ipynb

Big picture

Reinforcement learning is the study of sequential decision-making under uncertainty. This book develops the subject in seven stages:

  1. Foundations: agents, environments, Markov decision processes, and Bellman equations.
  2. Sample-based learning: Monte Carlo methods, temporal-difference learning, SARSA, and Q-learning.
  3. Exploration and stochastic approximation: bandits, regret, exploration strategies, and noisy recursive updates.
  4. Approximation: feature-based value functions and approximate dynamic programming.
  5. Policy optimization and deep RL: policy gradients, actor-critic methods, natural gradients, entropy regularization, DQN, PPO-style methods, and continuous control.
  6. Model-based and data-based RL: learned models, planning, offline RL, and imitation learning.
  7. Modern extensions and capstone: multi-agent RL, RLHF, and the final integrated project.

The book is designed so that classical mathematical foundations lead naturally to modern reinforcement-learning methods.

Conceptual dependency map

The main conceptual flow is:

RL loop
  -> Markov decision processes
  -> Bellman equations
  -> dynamic programming
  -> Monte Carlo and TD learning
  -> SARSA and Q-learning
  -> exploration and bandits
  -> stochastic approximation
  -> function approximation
  -> policy gradients and actor-critic methods
  -> deep RL and continuous control
  -> model-based RL and planning
  -> offline RL and imitation learning
  -> multi-agent RL and RLHF
  -> final integrated project

Another way to view the book is through four mathematical questions:

Question Main chapters
What is the decision problem? Chapters 1–3
How do we learn values from experience? Chapters 4–11
How do we optimize policies? Chapters 12–21
How do modern RL systems use models, data, demonstrations, games, and feedback? Chapters 22–28

Full chapter roadmap

Chapter Topic Main ideas Prerequisites Lab
1 Introduction to Reinforcement Learning RL loop, rewards, returns, policies, trajectories None Lab 1 · Colab
2 Markov Decision Processes Finite MDPs, transitions, policies, value functions Chapter 1 Lab 2 · Colab
3 Bellman Equations and Dynamic Programming Bellman equations, value iteration, policy iteration Chapters 1–2 Lab 3 · Colab
4 Monte Carlo Methods Sampled returns, first-visit/every-visit MC, MC control Chapters 1–3 Lab 4 · Colab
5 Temporal-Difference Learning TD error, TD(0), bootstrapping, TD control Chapters 1–4 Lab 5 · Colab
6 SARSA and Q-Learning in Depth On-policy/off-policy control, risk, cliff walking Chapter 5 Lab 6 · Colab
7 Exploration and Exploitation Epsilon-greedy, optimism, softmax, UCB Chapters 5–6 Lab 7 · Colab
8 Multi-Armed Bandits Regret, UCB, Thompson sampling, exploration theory Chapter 7 Lab 8 · Colab
9 Stochastic Approximation Robbins–Monro, step sizes, TD as stochastic approximation Chapters 5–8 Lab 9 · Colab
10 Function Approximation Features, regression viewpoint, approximation error Chapters 3–9 Lab 10 · Colab
11 Linear Value Function Approximation Projection, gradient MC, semi-gradient TD, LSTD Chapter 10 Lab 11 · Colab
12 Policy Gradient Methods Score-function identity, softmax policies, policy optimization Chapters 1–5 Lab 12 · Colab
13 REINFORCE Reward-to-go, baselines, variance reduction Chapter 12 Lab 13 · Colab
14 Actor-Critic Methods Actor, critic, TD advantage, bootstrapped policy gradients Chapters 12–13 Lab 14 · Colab
15 Natural Policy Gradient Fisher information, KL geometry, natural gradient Chapters 12–14 Lab 15 · Colab
16 Entropy-Regularized Reinforcement Learning Entropy bonuses, soft Bellman equations, temperature Chapters 12–15 Lab 16 · Colab
17 Approximate Dynamic Programming Fitted value iteration, rollout, approximate planning Chapters 3, 10–11 Lab 17 · Colab
18 Deep Q-Networks Neural Q-functions, replay buffers, target networks Chapters 10–11, 17 Lab 18 · Colab
19 Deep Policy Gradient Methods Neural policies, learned baselines, deep REINFORCE Chapters 12–14 Lab 19 · Colab
20 Trust Region and PPO-Style Methods Policy ratios, clipping, KL diagnostics Chapters 12–15, 19 Lab 20 · Colab
21 Continuous Control Gaussian policies, continuous actions, policy-gradient control Chapters 12–20 Lab 21 · Colab
22 Model-Based Reinforcement Learning Model estimation, planning in learned models, model error Chapters 2–3, 17 Lab 22 · Colab
23 Planning and Learning Dyna-Q, planning steps, prioritized sweeping, model staleness Chapter 22 Lab 23 · Colab
24 Offline Reinforcement Learning Fixed datasets, coverage, FQE, conservative offline FQI Chapters 5–6, 10–11 Lab 24 · Colab
25 Imitation Learning Expert demonstrations, behavior cloning, DAgger Chapters 2–6, 24 Lab 25 · Colab
26 Multi-Agent Reinforcement Learning Matrix games, Markov games, independent Q-learning Chapters 1–8 Lab 26 · Colab
27 Reinforcement Learning from Human Feedback Preferences, reward modeling, KL regularization, DPO-style learning Chapters 12–16, 24–25 Lab 27 · Colab
28 Final Integrated Reinforcement Learning Project Integrated comparison of DP, Q-learning, Dyna-Q, offline RL, imitation learning Most previous chapters Lab 28 · Colab

Module 1: Foundations

Chapters 1–3

The first module introduces the basic language of reinforcement learning.

Students learn:

  • agents and environments,
  • states and actions,
  • rewards and returns,
  • Markov decision processes,
  • policies,
  • value functions,
  • Bellman expectation equations,
  • Bellman optimality equations,
  • value iteration,
  • policy iteration.

The core equation is:

\[ V^*(s) = \max_a \left[ R(s,a) + \gamma \sum_{s'}P(s'\mid s,a)V^*(s') \right]. \]

Labs

  • Lab 1: Introduction to Reinforcement Learning
  • Lab 2: Markov Decision Processes
  • Lab 3: Bellman Equations and Dynamic Programming

Milestone

By the end of Module 1, students should be able to formulate a small finite MDP and solve it by dynamic programming.

Module 2: Learning from experience

Chapters 4–6

This module moves from known models to sampled experience.

Students learn:

  • Monte Carlo prediction,
  • Monte Carlo control,
  • temporal-difference prediction,
  • TD error,
  • bootstrapping,
  • SARSA,
  • Q-learning,
  • on-policy versus off-policy learning.

The central TD update is:

\[ V(S_t) \leftarrow V(S_t) + \alpha \left[ R_{t+1} + \gamma V(S_{t+1}) - V(S_t) \right]. \]

The central Q-learning update is:

\[ Q(S_t,A_t) \leftarrow Q(S_t,A_t) + \alpha \left[ R_{t+1} + \gamma \max_a Q(S_{t+1},a) - Q(S_t,A_t) \right]. \]

Labs

  • Lab 4: Monte Carlo Methods
  • Lab 5: Temporal-Difference Learning
  • Lab 6: SARSA and Q-Learning in Depth

Milestone

By the end of Module 2, students should be able to implement Monte Carlo prediction, TD prediction, SARSA, and Q-learning in a small environment.

Module 3: Exploration, bandits, and stochastic approximation

Chapters 7–9

This module studies how agents explore and how noisy updates behave.

Students learn:

  • exploration versus exploitation,
  • epsilon-greedy exploration,
  • optimistic initialization,
  • softmax action selection,
  • upper confidence bounds,
  • Thompson sampling,
  • regret,
  • Robbins–Monro stochastic approximation,
  • step-size schedules.

A basic epsilon-greedy rule is:

\[ A_t = \begin{cases} \text{random action}, & \text{with probability }\epsilon,\\ \arg\max_a Q(S_t,a), & \text{with probability }1-\epsilon. \end{cases} \]

Labs

  • Lab 7: Exploration and Exploitation
  • Lab 8: Multi-Armed Bandits
  • Lab 9: Stochastic Approximation

Milestone

By the end of Module 3, students should be able to compare exploration strategies using regret, learning curves, and state-action coverage.

Module 4: Function approximation

Chapters 10–11

This module introduces value approximation.

Students learn:

  • feature maps,
  • linear value functions,
  • regression viewpoint,
  • projection,
  • gradient Monte Carlo,
  • semi-gradient TD,
  • least-squares temporal difference ideas,
  • approximation error.

A linear value function is:

\[ V(s;w) = \phi(s)^T w. \]

Labs

  • Lab 10: Function Approximation
  • Lab 11: Linear Value Function Approximation

Milestone

By the end of Module 4, students should understand how tabular value functions generalize to parameterized value functions.

Module 5: Policy optimization

Chapters 12–16

This module studies direct policy optimization.

Students learn:

  • stochastic policies,
  • score-function identity,
  • REINFORCE,
  • reward-to-go,
  • baselines,
  • actor-critic methods,
  • TD advantage,
  • natural policy gradient,
  • Fisher information,
  • entropy regularization,
  • soft value functions.

The fundamental policy-gradient identity is:

\[ \nabla_\theta J(\theta) = E \left[ G_t \nabla_\theta \log \pi_\theta(A_t\mid S_t) \right]. \]

Entropy regularization adds an exploration term:

\[ J_\tau(\pi) = E_\pi \left[ \sum_t \gamma^t \left( R_{t+1} + \tau H(\pi(\cdot\mid S_t)) \right) \right]. \]

Labs

  • Lab 12: Policy Gradient Methods
  • Lab 13: REINFORCE
  • Lab 14: Actor-Critic Methods
  • Lab 15: Natural Policy Gradient
  • Lab 16: Entropy-Regularized Reinforcement Learning

Milestone

By the end of Module 5, students should be able to implement policy-gradient and actor-critic methods and explain the role of baselines, advantages, KL geometry, and entropy.

Module 6: Approximate and deep reinforcement learning

Chapters 17–21

This module connects approximate dynamic programming and policy optimization to deep RL.

Students learn:

  • fitted value iteration,
  • rollout improvement,
  • DQN-style learning,
  • replay buffers,
  • target networks,
  • neural policy gradients,
  • PPO-style clipping,
  • continuous control,
  • Gaussian policies.

A DQN-style target is:

\[ y = r + \gamma(1-d) \max_b Q(s',b;\theta^-). \]

A PPO-style ratio is:

\[ \rho_t(\theta) = \frac{\pi_\theta(A_t\mid S_t)}{\pi_{\text{old}}(A_t\mid S_t)}. \]

Labs

  • Lab 17: Approximate Dynamic Programming
  • Lab 18: Deep Q-Networks
  • Lab 19: Deep Policy Gradient Methods
  • Lab 20: Trust Region and PPO-Style Methods
  • Lab 21: Continuous Control

Milestone

By the end of Module 6, students should understand why function approximation changes reinforcement learning and why stabilizing tools such as replay buffers, target networks, clipping, and entropy matter.

Module 7: Model-based learning and planning

Chapters 22–23

This module studies agents that learn or use models.

Students learn:

  • estimating transition models,
  • estimating reward models,
  • planning in learned MDPs,
  • model error,
  • Dyna-Q,
  • simulated experience,
  • prioritized sweeping,
  • model staleness,
  • Dyna-Q+.

A learned model has the form:

\[ \widehat P(s'\mid s,a), \qquad \widehat R(s,a). \]

Dyna-Q combines real and simulated updates:

\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]

Labs

  • Lab 22: Model-Based Reinforcement Learning
  • Lab 23: Planning and Learning

Milestone

By the end of Module 7, students should be able to explain how planning reuses experience and how model error affects policy quality.

Module 8: Offline, imitation, multi-agent, and feedback-based RL

Chapters 24–27

This module introduces modern RL settings where the data source or learning objective changes.

Students learn:

  • offline RL from fixed datasets,
  • data coverage,
  • fitted Q evaluation,
  • conservative offline learning,
  • expert demonstrations,
  • behavior cloning,
  • DAgger,
  • normal-form games,
  • Markov games,
  • independent multi-agent learning,
  • preference data,
  • Bradley–Terry reward models,
  • KL-regularized policy optimization,
  • DPO-style preference learning.

Offline RL uses a fixed dataset:

\[ D = \{(s_i,a_i,r_i,s_i',d_i)\}_{i=1}^N. \]

RLHF preference modeling often uses:

\[ P(y_w \succ y_l\mid x) = \sigma \left( r_\theta(x,y_w)-r_\theta(x,y_l) \right). \]

Labs

  • Lab 24: Offline Reinforcement Learning
  • Lab 25: Imitation Learning
  • Lab 26: Multi-Agent Reinforcement Learning
  • Lab 27: Reinforcement Learning from Human Feedback

Milestone

By the end of Module 8, students should be able to reason about distribution shift, coverage, demonstrations, preference feedback, and interactions among multiple learning agents.

Module 9: Final integrated project

Chapter 28

The final chapter and lab integrate the book.

Students compare:

  • exact dynamic programming,
  • Q-learning,
  • Dyna-Q,
  • offline fitted Q iteration,
  • conservative offline learning,
  • behavior cloning.

The final project emphasizes:

  • common evaluation metrics,
  • learning curves,
  • policy visualizations,
  • coverage diagnostics,
  • policy disagreement analysis,
  • written explanation.

Lab

  • Lab 28: Final Integrated Reinforcement Learning Project

Milestone

By the end of the final project, students should be able to design and communicate a complete reinforcement-learning experiment.

Suggested semester schedule

A 14-week semester can use the following pacing.

Week Chapters Labs Theme
1 1–2 1–2 RL loop and MDPs
2 3 3 Bellman equations and dynamic programming
3 4–5 4–5 Monte Carlo and TD learning
4 6 6 SARSA and Q-learning
5 7–8 7–8 Exploration and bandits
6 9–10 9–10 Stochastic approximation and features
7 11–12 11–12 Linear value approximation and policy gradients
8 13–14 13–14 REINFORCE and actor-critic
9 15–16 15–16 Natural gradients and entropy regularization
10 17–18 17–18 Approximate DP and DQN
11 19–21 19–21 Deep policy gradients, PPO-style methods, continuous control
12 22–23 22–23 Model-based RL and planning
13 24–27 24–27 Offline RL, imitation learning, multi-agent RL, RLHF
14 28 28 Final integrated project

Accelerated 8-week path

For a shorter course or reading group:

Week Chapters Theme
1 1–3 MDPs and Bellman equations
2 4–6 Monte Carlo, TD, SARSA, Q-learning
3 7–9 Exploration, bandits, stochastic approximation
4 10–14 Function approximation and policy gradients
5 15–18 Natural gradients, entropy, ADP, DQN
6 19–23 Deep policy methods, continuous control, model-based RL
7 24–27 Offline RL, imitation learning, multi-agent RL, RLHF
8 28 Final project

Self-study path

For independent study, the following path is recommended.

Step 1: Build the foundation

Read Chapters 1–3 and complete Labs 1–3.

Goal:

Be able to define an MDP and solve a small one by value iteration.

Step 2: Learn from data

Read Chapters 4–6 and complete Labs 4–6.

Goal:

Be able to implement Monte Carlo, TD, SARSA, and Q-learning.

Step 3: Understand exploration

Read Chapters 7–9 and complete Labs 7–9.

Goal:

Be able to explain exploration-exploitation tradeoffs and compare bandit algorithms.

Step 4: Move beyond tables

Read Chapters 10–11 and complete Labs 10–11.

Goal:

Be able to use features to approximate value functions.

Step 5: Optimize policies directly

Read Chapters 12–16 and complete Labs 12–16.

Goal:

Be able to implement policy-gradient and actor-critic methods.

Step 6: Study modern RL algorithms

Read Chapters 17–23 and complete Labs 17–23.

Goal:

Be able to explain DQN, PPO-style updates, continuous control, model-based RL, and Dyna-Q.

Step 7: Study modern data and feedback settings

Read Chapters 24–27 and complete Labs 24–27.

Goal:

Be able to reason about offline data, demonstrations, multiple agents, and human preference feedback.

Step 8: Complete the final project

Complete Lab 28.

Goal:

Produce a short report comparing multiple RL approaches on one environment.

Mathematical skills developed

By following the roadmap, students develop the ability to:

  • manipulate Bellman equations,
  • reason about contraction mappings,
  • compute value and action-value functions,
  • analyze stochastic recursive updates,
  • understand exploration statistically,
  • interpret policy-gradient estimators,
  • connect KL divergence to policy stability,
  • diagnose approximation error,
  • reason about data distribution and coverage,
  • analyze simple game-theoretic learning systems,
  • model preferences using logistic likelihoods.

Computational skills developed

Students also learn to:

  • implement small environments,
  • simulate trajectories,
  • build tabular value and Q-functions,
  • write training loops,
  • evaluate policies fairly,
  • visualize values and policies,
  • compare learning curves,
  • create offline datasets,
  • train simple function approximators,
  • debug stochastic algorithms,
  • write reproducible notebooks.

Common choices for instructors

Classical RL path

Use Chapters 1–9 and Labs 1–9.

Best for a short introduction to MDPs, Bellman equations, MC, TD, Q-learning, exploration, and bandits.

Mathematical RL path

Use Chapters 1–16 and Labs 1–16.

Best for a mathematically focused course emphasizing Bellman operators, stochastic approximation, value approximation, and policy gradients.

Modern RL path

Use Chapters 10–28 and Labs 10–28 after a short MDP review.

Best for students who already know basic RL and want function approximation, deep RL, offline RL, imitation learning, multi-agent RL, and RLHF.

Project-centered path

Use Labs 1–6, 10–14, 18, 22–25, and 28.

Best for a course where the final project is central.

Common student difficulties

Difficulty 1: Confusing value and reward

Reward is immediate. Value is long-term.

\[ V^\pi(s) = E_\pi[G_t\mid S_t=s]. \]

Difficulty 2: Confusing prediction and control

Prediction evaluates a fixed policy. Control improves the policy.

Difficulty 3: Forgetting terminal-state handling

For terminal transitions, the target should usually not include a next-state value.

Difficulty 4: Comparing training and evaluation incorrectly

Training may use exploration. Evaluation usually uses the learned greedy or mean policy.

Difficulty 5: Ignoring random seeds

One run can be misleading. Use several seeds when possible.

Difficulty 6: Treating offline RL like online RL

Offline RL cannot collect new data. Coverage matters.

Difficulty 7: Over-optimizing learned rewards

In RLHF-style settings, optimizing a biased reward model too aggressively can cause reward hacking.

Final learning outcomes

After completing the roadmap, students should be able to:

  1. formulate sequential decision problems mathematically,
  2. derive Bellman equations,
  3. implement dynamic programming,
  4. learn values from sampled returns,
  5. implement TD learning and Q-learning,
  6. compare exploration strategies,
  7. understand stochastic approximation,
  8. use function approximation,
  9. implement policy-gradient methods,
  10. explain actor-critic learning,
  11. understand natural-gradient and entropy-regularized objectives,
  12. implement simplified deep RL ideas,
  13. use learned models for planning,
  14. diagnose offline RL failures,
  15. train imitation-learning policies,
  16. analyze multi-agent learning behavior,
  17. explain RLHF preference modeling,
  18. complete and communicate an integrated RL project.

Closing perspective

This roadmap is meant to show the unity of the subject.

The book begins with the Bellman equation, but the same ideas reappear throughout:

  • TD learning uses Bellman targets from samples.
  • Q-learning uses a Bellman optimality target.
  • Actor-critic methods use value estimates to improve policy gradients.
  • DQN uses neural networks to approximate Bellman fixed points.
  • Dyna-Q uses learned models to generate more Bellman updates.
  • Offline RL asks when Bellman updates are reliable with fixed data.
  • Imitation learning asks how demonstrations can define policies.
  • RLHF asks how preference feedback can define rewards.

The final goal is to see reinforcement learning as a mathematical language for prediction, control, planning, and learning under uncertainty.