Reinforcement Learning: A Mathematical Introduction
A mathematical and computational introduction with Python labs
The Mathematics of Reinforcement Learning
This book develops reinforcement learning from a mathematical point of view while keeping computation close to the theory. Each chapter is paired with a Google Colab-ready lab for independent study.
Book overview
The Mathematics of Reinforcement Learning is designed for students who want to understand reinforcement learning through the language of probability, optimization, dynamic programming, stochastic approximation, and statistical learning.
The book emphasizes four connected themes:
- Mathematical foundations: Markov decision processes, Bellman equations, value functions, policy gradients, stochastic approximation, and game-theoretic extensions.
- Algorithms: dynamic programming, Monte Carlo methods, temporal-difference learning, Q-learning, policy gradients, actor-critic methods, DQN, PPO-style updates, model-based RL, offline RL, imitation learning, multi-agent RL, and RLHF.
- Computation: every major idea is supported by Python examples and independent-study computer labs.
- Modern perspective: the later chapters connect classical RL theory to deep RL, offline learning, imitation learning, multi-agent learning, and reinforcement learning from human feedback.
Intended audience
This book is suitable for advanced undergraduate students, master’s students, and beginning doctoral students in applied mathematics, statistics, data science, machine learning, operations research, and related fields.
Recommended background:
- linear algebra,
- multivariable calculus,
- probability,
- basic optimization,
- basic Python programming.
Prior exposure to machine learning is helpful but not required.
Repository
The GitHub repository for this book is:
https://github.com/wanghemath/Book-MathRL
The computer labs are located in:
labs/
Each lab can be opened directly in Google Colab.
How to use this book
A good study path is:
- Read the chapter notes.
- Work through the corresponding lab.
- Modify the code examples.
- Complete the exercises.
- Write a short reflection explaining what the algorithm is optimizing and why it works.
- Use Lab 28 as a final integrated project.
Chapter and lab roadmap
| Chapter | Topic | Main ideas | Lab |
|---|---|---|---|
| 1 | Introduction to Reinforcement Learning | Agents, environments, rewards, returns, trajectories, and the RL loop. | Notebook · Open in Colab |
| 2 | Markov Decision Processes | Finite MDPs, transition kernels, rewards, policies, and value functions. | Notebook · Open in Colab |
| 3 | Bellman Equations and Dynamic Programming | Bellman expectation and optimality equations, value iteration, and policy iteration. | Notebook · Open in Colab |
| 4 | Monte Carlo Methods | Episode-based prediction and control from sampled returns. | Notebook · Open in Colab |
| 5 | Temporal-Difference Learning | TD prediction, bootstrapping, TD errors, and TD control. | Notebook · Open in Colab |
| 6 | SARSA and Q-Learning in Depth | On-policy and off-policy TD control, cliff walking, and risk under exploration. | Notebook · Open in Colab |
| 7 | Exploration and Exploitation | Epsilon-greedy, optimistic initialization, softmax exploration, and UCB. | Notebook · Open in Colab |
| 8 | Multi-Armed Bandits | Regret, UCB, Thompson sampling, and bandit experiments. | Notebook · Open in Colab |
| 9 | Stochastic Approximation | Robbins–Monro updates, step-size schedules, and TD as stochastic approximation. | Notebook · Open in Colab |
| 10 | Function Approximation | Feature maps, linear prediction, approximation error, and semi-gradient methods. | Notebook · Open in Colab |
| 11 | Linear Value Function Approximation | Projection, gradient Monte Carlo, semi-gradient TD, and LSTD ideas. | Notebook · Open in Colab |
| 12 | Policy Gradient Methods | Score-function identity, softmax policies, and direct policy optimization. | Notebook · Open in Colab |
| 13 | REINFORCE | Full-return and reward-to-go REINFORCE with variance reduction. | Notebook · Open in Colab |
| 14 | Actor-Critic Methods | Policy actors, value critics, TD advantages, and one-step actor-critic. | Notebook · Open in Colab |
| 15 | Natural Policy Gradient | Fisher geometry, KL-aware updates, and natural-gradient intuition. | Notebook · Open in Colab |
| 16 | Entropy-Regularized Reinforcement Learning | Soft value functions, entropy bonuses, temperature, and soft policy iteration. | Notebook · Open in Colab |
| 17 | Approximate Dynamic Programming | Fitted value iteration, rollout improvement, and approximation in planning. | Notebook · Open in Colab |
| 18 | Deep Q-Networks | Neural Q-functions, replay buffers, target networks, and DQN-style learning. | Notebook · Open in Colab |
| 19 | Deep Policy Gradient Methods | Neural stochastic policies, learned baselines, and deep REINFORCE-style training. | Notebook · Open in Colab |
| 20 | Trust Region and PPO-Style Methods | Policy ratios, KL diagnostics, clipping, and stable policy updates. | Notebook · Open in Colab |
| 21 | Continuous Control | Gaussian policies, continuous actions, and policy-gradient control. | Notebook · Open in Colab |
| 22 | Model-Based Reinforcement Learning | Learning transition/reward models and planning in estimated MDPs. | Notebook · Open in Colab |
| 23 | Planning and Learning | Dyna-Q, prioritized sweeping, model staleness, and Dyna-Q+. | Notebook · Open in Colab |
| 24 | Offline Reinforcement Learning | Fixed datasets, coverage, off-policy evaluation, and conservative offline learning. | Notebook · Open in Colab |
| 25 | Imitation Learning | Expert demonstrations, behavior cloning, distribution shift, and DAgger. | Notebook · Open in Colab |
| 26 | Multi-Agent Reinforcement Learning | Matrix games, Markov games, independent Q-learning, and coordination. | Notebook · Open in Colab |
| 27 | Reinforcement Learning from Human Feedback | Preference data, reward modeling, KL-regularized optimization, and DPO-style learning. | Notebook · Open in Colab |
| 28 | Final Integrated Reinforcement Learning Project | Capstone comparison of dynamic programming, Q-learning, Dyna-Q, offline RL, and imitation learning. | Notebook · Open in Colab |
Computer labs
The book includes 28 Google Colab-ready labs, one for each chapter.
The labs are designed for independent study. Each notebook includes:
- mathematical background before programming,
- complete Python implementations,
- visualizations,
- experiments,
- interpretation questions,
- exercises,
- a mini-project or final project component.
Example lab path:
labs/chapter-01-lab.ipynb
Example Colab URL pattern:
https://colab.research.google.com/github/wanghemath/Book-MathRL/blob/main/labs/chapter-01-lab.ipynb
Pedagogical structure
Each chapter follows a common pattern:
Conceptual motivation
Why the topic matters in reinforcement learning.Mathematical formulation
Definitions, equations, assumptions, and core theorems or principles.Algorithmic form
Pseudocode and computational interpretation.Python implementation
Small, readable examples that connect directly to the formulas.Experiments and diagnostics
Learning curves, policy visualizations, value plots, regret curves, or coverage plots.Exercises and projects
Questions that ask students to modify assumptions, compare algorithms, and explain results.
Major themes
Dynamic programming and Bellman equations
The first part of the book studies the mathematical core of RL:
\[ V^\pi(s) = \sum_a \pi(a\mid s) \left[ R(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V^\pi(s') \right], \]
and
\[ V^*(s) = \max_a \left[ R(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V^*(s') \right]. \]
Learning from samples
Monte Carlo and temporal-difference methods replace exact expectations by sampled experience:
\[ V(S_t) \leftarrow V(S_t) + \alpha \left[ R_{t+1}+\gamma V(S_{t+1})-V(S_t) \right]. \]
Control and exploration
Control methods learn how to choose actions, balancing exploitation with exploration:
\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]
Function approximation and deep RL
Modern RL uses parameterized value functions and policies:
\[ Q(s,a;\theta), \qquad V(s;w), \qquad \pi_\theta(a\mid s). \]
Policy optimization
Policy-gradient methods optimize policies directly:
\[ \nabla_\theta J(\theta) = E\left[ G_t\nabla_\theta \log \pi_\theta(A_t\mid S_t) \right]. \]
Modern extensions
The final part of the book introduces model-based RL, offline RL, imitation learning, multi-agent RL, and RLHF.
Final project
Lab 28 is the integrated capstone project. Students compare several RL approaches on one environment:
- exact dynamic programming,
- Q-learning,
- Dyna-Q,
- offline fitted Q iteration,
- conservative offline learning,
- behavior cloning.
The final report should include:
- an MDP formulation,
- algorithm descriptions,
- common evaluation metrics,
- learning curves,
- policy visualizations,
- coverage diagnostics,
- a concise explanation of why methods differ.
Suggested citation
If you use this book or its labs, please cite the repository:
He Wang, The Mathematics of Reinforcement Learning.
GitHub: https://github.com/wanghemath/Book-MathRL
License and use
This book is intended for educational use. Instructors may adapt the labs and examples for courses, workshops, and independent study, subject to the license terms of the repository.
Acknowledgment
This project is designed to help students see reinforcement learning not only as a collection of algorithms, but as a coherent mathematical language for sequential decision-making under uncertainty.