Roadmap
A learning path through The Mathematics of Reinforcement Learning
Roadmap
This roadmap explains how the chapters, labs, mathematical ideas, and projects fit together. It can be used for a semester course, an accelerated module, or independent study.
Repository
The GitHub repository for the book is:
https://github.com/wanghemath/Book-MathRL
The labs are stored in:
labs/
The final integrated project is:
labs/chapter-28-lab.ipynb
Big picture
Reinforcement learning is the study of sequential decision-making under uncertainty. This book develops the subject in seven stages:
- Foundations: agents, environments, Markov decision processes, and Bellman equations.
- Sample-based learning: Monte Carlo methods, temporal-difference learning, SARSA, and Q-learning.
- Exploration and stochastic approximation: bandits, regret, exploration strategies, and noisy recursive updates.
- Approximation: feature-based value functions and approximate dynamic programming.
- Policy optimization and deep RL: policy gradients, actor-critic methods, natural gradients, entropy regularization, DQN, PPO-style methods, and continuous control.
- Model-based and data-based RL: learned models, planning, offline RL, and imitation learning.
- Modern extensions and capstone: multi-agent RL, RLHF, and the final integrated project.
The book is designed so that classical mathematical foundations lead naturally to modern reinforcement-learning methods.
Conceptual dependency map
The main conceptual flow is:
RL loop
-> Markov decision processes
-> Bellman equations
-> dynamic programming
-> Monte Carlo and TD learning
-> SARSA and Q-learning
-> exploration and bandits
-> stochastic approximation
-> function approximation
-> policy gradients and actor-critic methods
-> deep RL and continuous control
-> model-based RL and planning
-> offline RL and imitation learning
-> multi-agent RL and RLHF
-> final integrated project
Another way to view the book is through four mathematical questions:
| Question | Main chapters |
|---|---|
| What is the decision problem? | Chapters 1–3 |
| How do we learn values from experience? | Chapters 4–11 |
| How do we optimize policies? | Chapters 12–21 |
| How do modern RL systems use models, data, demonstrations, games, and feedback? | Chapters 22–28 |
Full chapter roadmap
| Chapter | Topic | Main ideas | Prerequisites | Lab |
|---|---|---|---|---|
| 1 | Introduction to Reinforcement Learning | RL loop, rewards, returns, policies, trajectories | None | Lab 1 · Colab |
| 2 | Markov Decision Processes | Finite MDPs, transitions, policies, value functions | Chapter 1 | Lab 2 · Colab |
| 3 | Bellman Equations and Dynamic Programming | Bellman equations, value iteration, policy iteration | Chapters 1–2 | Lab 3 · Colab |
| 4 | Monte Carlo Methods | Sampled returns, first-visit/every-visit MC, MC control | Chapters 1–3 | Lab 4 · Colab |
| 5 | Temporal-Difference Learning | TD error, TD(0), bootstrapping, TD control | Chapters 1–4 | Lab 5 · Colab |
| 6 | SARSA and Q-Learning in Depth | On-policy/off-policy control, risk, cliff walking | Chapter 5 | Lab 6 · Colab |
| 7 | Exploration and Exploitation | Epsilon-greedy, optimism, softmax, UCB | Chapters 5–6 | Lab 7 · Colab |
| 8 | Multi-Armed Bandits | Regret, UCB, Thompson sampling, exploration theory | Chapter 7 | Lab 8 · Colab |
| 9 | Stochastic Approximation | Robbins–Monro, step sizes, TD as stochastic approximation | Chapters 5–8 | Lab 9 · Colab |
| 10 | Function Approximation | Features, regression viewpoint, approximation error | Chapters 3–9 | Lab 10 · Colab |
| 11 | Linear Value Function Approximation | Projection, gradient MC, semi-gradient TD, LSTD | Chapter 10 | Lab 11 · Colab |
| 12 | Policy Gradient Methods | Score-function identity, softmax policies, policy optimization | Chapters 1–5 | Lab 12 · Colab |
| 13 | REINFORCE | Reward-to-go, baselines, variance reduction | Chapter 12 | Lab 13 · Colab |
| 14 | Actor-Critic Methods | Actor, critic, TD advantage, bootstrapped policy gradients | Chapters 12–13 | Lab 14 · Colab |
| 15 | Natural Policy Gradient | Fisher information, KL geometry, natural gradient | Chapters 12–14 | Lab 15 · Colab |
| 16 | Entropy-Regularized Reinforcement Learning | Entropy bonuses, soft Bellman equations, temperature | Chapters 12–15 | Lab 16 · Colab |
| 17 | Approximate Dynamic Programming | Fitted value iteration, rollout, approximate planning | Chapters 3, 10–11 | Lab 17 · Colab |
| 18 | Deep Q-Networks | Neural Q-functions, replay buffers, target networks | Chapters 10–11, 17 | Lab 18 · Colab |
| 19 | Deep Policy Gradient Methods | Neural policies, learned baselines, deep REINFORCE | Chapters 12–14 | Lab 19 · Colab |
| 20 | Trust Region and PPO-Style Methods | Policy ratios, clipping, KL diagnostics | Chapters 12–15, 19 | Lab 20 · Colab |
| 21 | Continuous Control | Gaussian policies, continuous actions, policy-gradient control | Chapters 12–20 | Lab 21 · Colab |
| 22 | Model-Based Reinforcement Learning | Model estimation, planning in learned models, model error | Chapters 2–3, 17 | Lab 22 · Colab |
| 23 | Planning and Learning | Dyna-Q, planning steps, prioritized sweeping, model staleness | Chapter 22 | Lab 23 · Colab |
| 24 | Offline Reinforcement Learning | Fixed datasets, coverage, FQE, conservative offline FQI | Chapters 5–6, 10–11 | Lab 24 · Colab |
| 25 | Imitation Learning | Expert demonstrations, behavior cloning, DAgger | Chapters 2–6, 24 | Lab 25 · Colab |
| 26 | Multi-Agent Reinforcement Learning | Matrix games, Markov games, independent Q-learning | Chapters 1–8 | Lab 26 · Colab |
| 27 | Reinforcement Learning from Human Feedback | Preferences, reward modeling, KL regularization, DPO-style learning | Chapters 12–16, 24–25 | Lab 27 · Colab |
| 28 | Final Integrated Reinforcement Learning Project | Integrated comparison of DP, Q-learning, Dyna-Q, offline RL, imitation learning | Most previous chapters | Lab 28 · Colab |
Module 1: Foundations
Chapters 1–3
The first module introduces the basic language of reinforcement learning.
Students learn:
- agents and environments,
- states and actions,
- rewards and returns,
- Markov decision processes,
- policies,
- value functions,
- Bellman expectation equations,
- Bellman optimality equations,
- value iteration,
- policy iteration.
The core equation is:
\[ V^*(s) = \max_a \left[ R(s,a) + \gamma \sum_{s'}P(s'\mid s,a)V^*(s') \right]. \]
Labs
- Lab 1: Introduction to Reinforcement Learning
- Lab 2: Markov Decision Processes
- Lab 3: Bellman Equations and Dynamic Programming
Milestone
By the end of Module 1, students should be able to formulate a small finite MDP and solve it by dynamic programming.
Module 2: Learning from experience
Chapters 4–6
This module moves from known models to sampled experience.
Students learn:
- Monte Carlo prediction,
- Monte Carlo control,
- temporal-difference prediction,
- TD error,
- bootstrapping,
- SARSA,
- Q-learning,
- on-policy versus off-policy learning.
The central TD update is:
\[ V(S_t) \leftarrow V(S_t) + \alpha \left[ R_{t+1} + \gamma V(S_{t+1}) - V(S_t) \right]. \]
The central Q-learning update is:
\[ Q(S_t,A_t) \leftarrow Q(S_t,A_t) + \alpha \left[ R_{t+1} + \gamma \max_a Q(S_{t+1},a) - Q(S_t,A_t) \right]. \]
Labs
- Lab 4: Monte Carlo Methods
- Lab 5: Temporal-Difference Learning
- Lab 6: SARSA and Q-Learning in Depth
Milestone
By the end of Module 2, students should be able to implement Monte Carlo prediction, TD prediction, SARSA, and Q-learning in a small environment.
Module 3: Exploration, bandits, and stochastic approximation
Chapters 7–9
This module studies how agents explore and how noisy updates behave.
Students learn:
- exploration versus exploitation,
- epsilon-greedy exploration,
- optimistic initialization,
- softmax action selection,
- upper confidence bounds,
- Thompson sampling,
- regret,
- Robbins–Monro stochastic approximation,
- step-size schedules.
A basic epsilon-greedy rule is:
\[ A_t = \begin{cases} \text{random action}, & \text{with probability }\epsilon,\\ \arg\max_a Q(S_t,a), & \text{with probability }1-\epsilon. \end{cases} \]
Labs
- Lab 7: Exploration and Exploitation
- Lab 8: Multi-Armed Bandits
- Lab 9: Stochastic Approximation
Milestone
By the end of Module 3, students should be able to compare exploration strategies using regret, learning curves, and state-action coverage.
Module 4: Function approximation
Chapters 10–11
This module introduces value approximation.
Students learn:
- feature maps,
- linear value functions,
- regression viewpoint,
- projection,
- gradient Monte Carlo,
- semi-gradient TD,
- least-squares temporal difference ideas,
- approximation error.
A linear value function is:
\[ V(s;w) = \phi(s)^T w. \]
Labs
- Lab 10: Function Approximation
- Lab 11: Linear Value Function Approximation
Milestone
By the end of Module 4, students should understand how tabular value functions generalize to parameterized value functions.
Module 5: Policy optimization
Chapters 12–16
This module studies direct policy optimization.
Students learn:
- stochastic policies,
- score-function identity,
- REINFORCE,
- reward-to-go,
- baselines,
- actor-critic methods,
- TD advantage,
- natural policy gradient,
- Fisher information,
- entropy regularization,
- soft value functions.
The fundamental policy-gradient identity is:
\[ \nabla_\theta J(\theta) = E \left[ G_t \nabla_\theta \log \pi_\theta(A_t\mid S_t) \right]. \]
Entropy regularization adds an exploration term:
\[ J_\tau(\pi) = E_\pi \left[ \sum_t \gamma^t \left( R_{t+1} + \tau H(\pi(\cdot\mid S_t)) \right) \right]. \]
Labs
- Lab 12: Policy Gradient Methods
- Lab 13: REINFORCE
- Lab 14: Actor-Critic Methods
- Lab 15: Natural Policy Gradient
- Lab 16: Entropy-Regularized Reinforcement Learning
Milestone
By the end of Module 5, students should be able to implement policy-gradient and actor-critic methods and explain the role of baselines, advantages, KL geometry, and entropy.
Module 6: Approximate and deep reinforcement learning
Chapters 17–21
This module connects approximate dynamic programming and policy optimization to deep RL.
Students learn:
- fitted value iteration,
- rollout improvement,
- DQN-style learning,
- replay buffers,
- target networks,
- neural policy gradients,
- PPO-style clipping,
- continuous control,
- Gaussian policies.
A DQN-style target is:
\[ y = r + \gamma(1-d) \max_b Q(s',b;\theta^-). \]
A PPO-style ratio is:
\[ \rho_t(\theta) = \frac{\pi_\theta(A_t\mid S_t)}{\pi_{\text{old}}(A_t\mid S_t)}. \]
Labs
- Lab 17: Approximate Dynamic Programming
- Lab 18: Deep Q-Networks
- Lab 19: Deep Policy Gradient Methods
- Lab 20: Trust Region and PPO-Style Methods
- Lab 21: Continuous Control
Milestone
By the end of Module 6, students should understand why function approximation changes reinforcement learning and why stabilizing tools such as replay buffers, target networks, clipping, and entropy matter.
Module 7: Model-based learning and planning
Chapters 22–23
This module studies agents that learn or use models.
Students learn:
- estimating transition models,
- estimating reward models,
- planning in learned MDPs,
- model error,
- Dyna-Q,
- simulated experience,
- prioritized sweeping,
- model staleness,
- Dyna-Q+.
A learned model has the form:
\[ \widehat P(s'\mid s,a), \qquad \widehat R(s,a). \]
Dyna-Q combines real and simulated updates:
\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]
Labs
- Lab 22: Model-Based Reinforcement Learning
- Lab 23: Planning and Learning
Milestone
By the end of Module 7, students should be able to explain how planning reuses experience and how model error affects policy quality.
Module 8: Offline, imitation, multi-agent, and feedback-based RL
Chapters 24–27
This module introduces modern RL settings where the data source or learning objective changes.
Students learn:
- offline RL from fixed datasets,
- data coverage,
- fitted Q evaluation,
- conservative offline learning,
- expert demonstrations,
- behavior cloning,
- DAgger,
- normal-form games,
- Markov games,
- independent multi-agent learning,
- preference data,
- Bradley–Terry reward models,
- KL-regularized policy optimization,
- DPO-style preference learning.
Offline RL uses a fixed dataset:
\[ D = \{(s_i,a_i,r_i,s_i',d_i)\}_{i=1}^N. \]
RLHF preference modeling often uses:
\[ P(y_w \succ y_l\mid x) = \sigma \left( r_\theta(x,y_w)-r_\theta(x,y_l) \right). \]
Labs
- Lab 24: Offline Reinforcement Learning
- Lab 25: Imitation Learning
- Lab 26: Multi-Agent Reinforcement Learning
- Lab 27: Reinforcement Learning from Human Feedback
Milestone
By the end of Module 8, students should be able to reason about distribution shift, coverage, demonstrations, preference feedback, and interactions among multiple learning agents.
Module 9: Final integrated project
Chapter 28
The final chapter and lab integrate the book.
Students compare:
- exact dynamic programming,
- Q-learning,
- Dyna-Q,
- offline fitted Q iteration,
- conservative offline learning,
- behavior cloning.
The final project emphasizes:
- common evaluation metrics,
- learning curves,
- policy visualizations,
- coverage diagnostics,
- policy disagreement analysis,
- written explanation.
Lab
- Lab 28: Final Integrated Reinforcement Learning Project
Milestone
By the end of the final project, students should be able to design and communicate a complete reinforcement-learning experiment.
Suggested semester schedule
A 14-week semester can use the following pacing.
| Week | Chapters | Labs | Theme |
|---|---|---|---|
| 1 | 1–2 | 1–2 | RL loop and MDPs |
| 2 | 3 | 3 | Bellman equations and dynamic programming |
| 3 | 4–5 | 4–5 | Monte Carlo and TD learning |
| 4 | 6 | 6 | SARSA and Q-learning |
| 5 | 7–8 | 7–8 | Exploration and bandits |
| 6 | 9–10 | 9–10 | Stochastic approximation and features |
| 7 | 11–12 | 11–12 | Linear value approximation and policy gradients |
| 8 | 13–14 | 13–14 | REINFORCE and actor-critic |
| 9 | 15–16 | 15–16 | Natural gradients and entropy regularization |
| 10 | 17–18 | 17–18 | Approximate DP and DQN |
| 11 | 19–21 | 19–21 | Deep policy gradients, PPO-style methods, continuous control |
| 12 | 22–23 | 22–23 | Model-based RL and planning |
| 13 | 24–27 | 24–27 | Offline RL, imitation learning, multi-agent RL, RLHF |
| 14 | 28 | 28 | Final integrated project |
Accelerated 8-week path
For a shorter course or reading group:
| Week | Chapters | Theme |
|---|---|---|
| 1 | 1–3 | MDPs and Bellman equations |
| 2 | 4–6 | Monte Carlo, TD, SARSA, Q-learning |
| 3 | 7–9 | Exploration, bandits, stochastic approximation |
| 4 | 10–14 | Function approximation and policy gradients |
| 5 | 15–18 | Natural gradients, entropy, ADP, DQN |
| 6 | 19–23 | Deep policy methods, continuous control, model-based RL |
| 7 | 24–27 | Offline RL, imitation learning, multi-agent RL, RLHF |
| 8 | 28 | Final project |
Self-study path
For independent study, the following path is recommended.
Step 1: Build the foundation
Read Chapters 1–3 and complete Labs 1–3.
Goal:
Be able to define an MDP and solve a small one by value iteration.
Step 2: Learn from data
Read Chapters 4–6 and complete Labs 4–6.
Goal:
Be able to implement Monte Carlo, TD, SARSA, and Q-learning.
Step 3: Understand exploration
Read Chapters 7–9 and complete Labs 7–9.
Goal:
Be able to explain exploration-exploitation tradeoffs and compare bandit algorithms.
Step 4: Move beyond tables
Read Chapters 10–11 and complete Labs 10–11.
Goal:
Be able to use features to approximate value functions.
Step 5: Optimize policies directly
Read Chapters 12–16 and complete Labs 12–16.
Goal:
Be able to implement policy-gradient and actor-critic methods.
Step 6: Study modern RL algorithms
Read Chapters 17–23 and complete Labs 17–23.
Goal:
Be able to explain DQN, PPO-style updates, continuous control, model-based RL, and Dyna-Q.
Step 7: Study modern data and feedback settings
Read Chapters 24–27 and complete Labs 24–27.
Goal:
Be able to reason about offline data, demonstrations, multiple agents, and human preference feedback.
Step 8: Complete the final project
Complete Lab 28.
Goal:
Produce a short report comparing multiple RL approaches on one environment.
Recommended project checkpoints
For a course project, use the following checkpoints.
| Checkpoint | Deliverable |
|---|---|
| Proposal | Project question, environment, methods, metrics |
| Baseline | Working environment and one baseline algorithm |
| Comparison | At least two methods compared under common metrics |
| Diagnostics | Learning curve, policy plot, coverage or error analysis |
| Draft report | Initial write-up with equations and results |
| Final report | Complete project with conclusions and reproducible notebook |
Mathematical skills developed
By following the roadmap, students develop the ability to:
- manipulate Bellman equations,
- reason about contraction mappings,
- compute value and action-value functions,
- analyze stochastic recursive updates,
- understand exploration statistically,
- interpret policy-gradient estimators,
- connect KL divergence to policy stability,
- diagnose approximation error,
- reason about data distribution and coverage,
- analyze simple game-theoretic learning systems,
- model preferences using logistic likelihoods.
Computational skills developed
Students also learn to:
- implement small environments,
- simulate trajectories,
- build tabular value and Q-functions,
- write training loops,
- evaluate policies fairly,
- visualize values and policies,
- compare learning curves,
- create offline datasets,
- train simple function approximators,
- debug stochastic algorithms,
- write reproducible notebooks.
Common choices for instructors
Classical RL path
Use Chapters 1–9 and Labs 1–9.
Best for a short introduction to MDPs, Bellman equations, MC, TD, Q-learning, exploration, and bandits.
Mathematical RL path
Use Chapters 1–16 and Labs 1–16.
Best for a mathematically focused course emphasizing Bellman operators, stochastic approximation, value approximation, and policy gradients.
Modern RL path
Use Chapters 10–28 and Labs 10–28 after a short MDP review.
Best for students who already know basic RL and want function approximation, deep RL, offline RL, imitation learning, multi-agent RL, and RLHF.
Project-centered path
Use Labs 1–6, 10–14, 18, 22–25, and 28.
Best for a course where the final project is central.
Common student difficulties
Difficulty 1: Confusing value and reward
Reward is immediate. Value is long-term.
\[ V^\pi(s) = E_\pi[G_t\mid S_t=s]. \]
Difficulty 2: Confusing prediction and control
Prediction evaluates a fixed policy. Control improves the policy.
Difficulty 3: Forgetting terminal-state handling
For terminal transitions, the target should usually not include a next-state value.
Difficulty 4: Comparing training and evaluation incorrectly
Training may use exploration. Evaluation usually uses the learned greedy or mean policy.
Difficulty 5: Ignoring random seeds
One run can be misleading. Use several seeds when possible.
Difficulty 6: Treating offline RL like online RL
Offline RL cannot collect new data. Coverage matters.
Difficulty 7: Over-optimizing learned rewards
In RLHF-style settings, optimizing a biased reward model too aggressively can cause reward hacking.
Final learning outcomes
After completing the roadmap, students should be able to:
- formulate sequential decision problems mathematically,
- derive Bellman equations,
- implement dynamic programming,
- learn values from sampled returns,
- implement TD learning and Q-learning,
- compare exploration strategies,
- understand stochastic approximation,
- use function approximation,
- implement policy-gradient methods,
- explain actor-critic learning,
- understand natural-gradient and entropy-regularized objectives,
- implement simplified deep RL ideas,
- use learned models for planning,
- diagnose offline RL failures,
- train imitation-learning policies,
- analyze multi-agent learning behavior,
- explain RLHF preference modeling,
- complete and communicate an integrated RL project.
Closing perspective
This roadmap is meant to show the unity of the subject.
The book begins with the Bellman equation, but the same ideas reappear throughout:
- TD learning uses Bellman targets from samples.
- Q-learning uses a Bellman optimality target.
- Actor-critic methods use value estimates to improve policy gradients.
- DQN uses neural networks to approximate Bellman fixed points.
- Dyna-Q uses learned models to generate more Bellman updates.
- Offline RL asks when Bellman updates are reliable with fixed data.
- Imitation learning asks how demonstrations can define policies.
- RLHF asks how preference feedback can define rewards.
The final goal is to see reinforcement learning as a mathematical language for prediction, control, planning, and learning under uncertainty.