Introduction
Why reinforcement learning is a mathematical language for sequential decision-making
Introduction
Reinforcement learning studies how an agent learns to make decisions through interaction with an environment. At each time step, the agent observes a state, chooses an action, receives a reward, and moves to a new state. Over time, the agent tries to improve its behavior.
This simple loop contains a rich mathematical structure:
\[ S_t \longrightarrow A_t \longrightarrow R_{t+1}, S_{t+1}. \]
The central question is:
How should an agent choose actions now in order to obtain good long-term outcomes later?
This book develops reinforcement learning as a mathematical subject supported by computation. The goal is not only to implement algorithms, but to understand the equations, assumptions, approximations, and tradeoffs behind them.
Why reinforcement learning?
Many problems are not one-shot prediction problems. They are sequential decision problems.
Examples include:
- a robot learning to move,
- a student-learning platform choosing the next exercise,
- a portfolio strategy adjusting over time,
- a recommender system adapting to feedback,
- a game-playing agent planning several moves ahead,
- a language-model assistant improving responses from preference feedback,
- a medical decision system balancing short-term and long-term outcomes.
In each case, actions influence future states. Therefore, the quality of an action cannot be judged only by its immediate reward. It must be judged by its future consequences.
This is the reason reinforcement learning needs mathematics beyond ordinary supervised learning.
The reinforcement-learning loop
A standard reinforcement-learning problem has five main ingredients:
| Object | Meaning |
|---|---|
| \(S_t\) | state at time \(t\) |
| \(A_t\) | action chosen by the agent |
| \(R_{t+1}\) | reward received after acting |
| \(S_{t+1}\) | next state |
| \(\pi(a\mid s)\) | policy, or rule for choosing actions |
The agent’s goal is usually to maximize expected discounted return:
\[ G_t = R_{t+1} + \gamma R_{t+2} + \gamma^2 R_{t+3} + \cdots, \]
where
\[ 0\le \gamma < 1 \]
is the discount factor.
The discount factor controls how much the agent values future rewards.
- If \(\gamma\) is small, the agent is short-sighted.
- If \(\gamma\) is close to \(1\), the agent is far-sighted.
The mathematical core
The mathematical foundation of reinforcement learning is the Markov decision process, or MDP.
An MDP consists of:
- a state space \(S\),
- an action space \(A\),
- transition probabilities \(P(s'\mid s,a)\),
- rewards \(R(s,a)\) or \(R(s,a,s')\),
- a discount factor \(\gamma\).
A policy \(\pi\) induces a value function:
\[ V^\pi(s) = E_\pi \left[ G_t \mid S_t=s \right]. \]
The corresponding action-value function is
\[ Q^\pi(s,a) = E_\pi \left[ G_t \mid S_t=s,\ A_t=a \right]. \]
The key recursive structure is the Bellman equation:
\[ V^\pi(s) = \sum_a \pi(a\mid s) \left[ R(s,a) + \gamma \sum_{s'} P(s'\mid s,a)V^\pi(s') \right]. \]
The optimal value function satisfies the Bellman optimality equation:
\[ V^*(s) = \max_a \left[ R(s,a) + \gamma \sum_{s'}P(s'\mid s,a)V^*(s') \right]. \]
These equations are the backbone of the book.
From exact mathematics to learning algorithms
If the transition probabilities and rewards are known, we can solve the MDP by dynamic programming.
But in many problems, the model is unknown. The agent must learn from sampled experience.
This leads to Monte Carlo and temporal-difference methods.
A Monte Carlo method estimates value from complete sampled returns:
\[ V(s) \approx \text{average of returns observed after visiting }s. \]
A temporal-difference method updates values before the episode is finished:
\[ V(S_t) \leftarrow V(S_t) + \alpha \left[ R_{t+1} + \gamma V(S_{t+1}) - V(S_t) \right]. \]
The quantity
\[ \delta_t = R_{t+1} + \gamma V(S_{t+1}) - V(S_t) \]
is called the TD error.
It measures the surprise in the value estimate after seeing one transition.
Control: learning how to act
Prediction estimates values for a fixed policy. Control improves the policy.
A central control algorithm is Q-learning:
\[ Q(S_t,A_t) \leftarrow Q(S_t,A_t) + \alpha \left[ R_{t+1} + \gamma\max_a Q(S_{t+1},a) - Q(S_t,A_t) \right]. \]
This update learns action values directly. After learning, the agent can act greedily:
\[ \pi(s) = \arg\max_a Q(s,a). \]
But control introduces a new problem: exploration.
The agent must sometimes try actions that do not currently look best. Otherwise, it may never discover better long-term strategies.
Exploration and exploitation
Reinforcement learning balances two forces:
- exploitation: choose actions that currently look best,
- exploration: try uncertain actions to gain information.
A simple exploration strategy is epsilon-greedy:
\[ A_t = \begin{cases} \text{random action}, & \text{with probability }\epsilon,\\ \arg\max_a Q(S_t,a), & \text{with probability }1-\epsilon. \end{cases} \]
Later chapters study more refined exploration ideas, including optimistic initialization, softmax exploration, upper confidence bounds, and Thompson sampling.
Function approximation
Tabular methods store one value for each state or state-action pair. This becomes impossible when the state space is large or continuous.
Function approximation replaces tables by parameterized functions:
\[ V(s;w), \qquad Q(s,a;\theta), \qquad \pi_\theta(a\mid s). \]
A linear value function has the form
\[ V(s;w) = \phi(s)^T w, \]
where \(\phi(s)\) is a feature vector.
Deep reinforcement learning uses neural networks, for example:
\[ Q(s,a;\theta) \approx Q^*(s,a). \]
This book introduces function approximation carefully before moving to deep Q-networks and deep policy-gradient methods.
Policy gradients
Value-based methods learn values first and then choose actions from values.
Policy-gradient methods optimize the policy directly.
A stochastic policy has parameters \(\theta\):
\[ \pi_\theta(a\mid s). \]
The objective is
\[ J(\theta) = E_{\pi_\theta}[G_0]. \]
A fundamental policy-gradient identity is
\[ \nabla_\theta J(\theta) = E_{\pi_\theta} \left[ G_t \nabla_\theta \log \pi_\theta(A_t\mid S_t) \right]. \]
This identity leads to REINFORCE, actor-critic methods, natural policy gradients, PPO-style methods, and continuous-control algorithms.
Modern reinforcement learning topics
The later chapters connect classical reinforcement-learning theory to modern directions.
Deep reinforcement learning
Deep RL combines reinforcement learning with neural networks. The book introduces DQN, replay buffers, target networks, deep policy gradients, and PPO-style clipping.
Model-based reinforcement learning
Model-based RL estimates or uses a model:
\[ \widehat P(s'\mid s,a), \qquad \widehat R(s,a), \]
then plans inside the learned model.
Planning and learning
Dyna-Q combines real experience with simulated model-based updates. This connects reinforcement learning to planning, search, and model predictive control.
Offline reinforcement learning
Offline RL learns from a fixed dataset. It is important when new interaction is expensive or risky. The main mathematical challenge is distribution shift and insufficient coverage.
Imitation learning
Imitation learning trains policies from expert demonstrations:
\[ D_E=\{(s_i,a_i)\}_{i=1}^N. \]
Behavior cloning treats policy learning as supervised classification, while DAgger-style methods address distribution shift.
Multi-agent reinforcement learning
Multi-agent RL studies several agents acting in the same environment. The book introduces matrix games, Markov games, coordination, competition, and independent Q-learning.
Reinforcement learning from human feedback
RLHF uses preference data to train reward models and optimize policies. The book gives a mathematical introduction to preference modeling, KL regularization, and DPO-style objectives.
Computational philosophy
The computer labs are not separate from the mathematics. They are part of the mathematical learning process.
Each lab asks students to:
- read the mathematical setup,
- implement the algorithm,
- visualize the result,
- change a parameter,
- explain what changed and why.
The goal is to develop computational intuition for equations such as:
\[ V^\pi = T^\pi V^\pi, \]
\[ V^* = T^*V^*, \]
\[ Q \leftarrow Q+\alpha(\text{target}-Q), \]
and
\[ \nabla_\theta J(\theta) = E[ \text{advantage} \times \nabla_\theta\log\pi_\theta(A_t\mid S_t) ]. \]
Python and Google Colab labs
The repository for the book is:
https://github.com/wanghemath/Book-MathRL
The labs are stored in:
labs/
Every lab is designed to run in Google Colab.
For example:
labs/chapter-01-lab.ipynb
can be opened using:
https://colab.research.google.com/github/wanghemath/Book-MathRL/blob/main/labs/chapter-01-lab.ipynb
The labs use lightweight dependencies whenever possible:
- NumPy,
- pandas,
- matplotlib.
Some labs implement neural-network ideas manually so that students can see the mathematics of the gradients.
Suggested study plan
A complete path through the book is:
- Chapters 1–3: learn the MDP and Bellman-equation foundations.
- Chapters 4–6: learn sample-based prediction and control.
- Chapters 7–9: study exploration, bandits, and stochastic approximation.
- Chapters 10–11: understand function approximation.
- Chapters 12–16: study policy gradients, actor-critic methods, natural gradients, and entropy regularization.
- Chapters 17–21: move toward approximate and deep reinforcement learning.
- Chapters 22–23: study model-based RL and planning.
- Chapters 24–27: study modern topics: offline RL, imitation learning, multi-agent RL, and RLHF.
- Chapter 28: complete the final integrated project.
How to read mathematical formulas
This book uses standard notation, but several symbols appear repeatedly.
| Symbol | Meaning |
|---|---|
| \(S_t\) | state at time \(t\) |
| \(A_t\) | action at time \(t\) |
| \(R_{t+1}\) | reward after taking action \(A_t\) |
| \(\gamma\) | discount factor |
| \(\pi\) | policy |
| \(V^\pi\) | value function of policy \(\pi\) |
| \(Q^\pi\) | action-value function of policy \(\pi\) |
| \(V^*\) | optimal value function |
| \(Q^*\) | optimal action-value function |
| \(\alpha\) | learning rate |
| \(\epsilon\) | exploration probability |
| \(\delta_t\) | temporal-difference error |
| \(\theta,w\) | model or policy parameters |
What students should be able to do by the end
After completing the book and labs, students should be able to:
- formulate a sequential decision problem as an MDP,
- derive Bellman equations for a given policy,
- implement value iteration and policy iteration,
- implement Monte Carlo and TD learning,
- compare SARSA and Q-learning,
- explain exploration-exploitation tradeoffs,
- use bandit algorithms and regret curves,
- understand stochastic approximation in RL,
- build simple function approximators,
- implement policy-gradient and actor-critic methods,
- explain entropy regularization and soft value functions,
- implement DQN-style and PPO-style ideas in small examples,
- reason about continuous control,
- estimate models and plan with them,
- diagnose offline RL coverage problems,
- train imitation-learning policies from demonstrations,
- analyze simple multi-agent RL systems,
- explain the mathematical structure of RLHF,
- complete an integrated RL project and write a clear report.
A note on rigor and computation
Reinforcement learning sits at the intersection of probability, optimization, statistics, control, and computation. A rigorous understanding requires both symbolic reasoning and numerical experimentation.
This book therefore treats code as a companion to proof and derivation.
When studying an algorithm, ask:
- What objective is it trying to optimize?
- What equation or fixed point defines the target?
- What approximation is being made?
- What data distribution is used?
- What can go wrong?
- How can we diagnose failure?
These questions are often more important than memorizing the algorithm.
Final project
The final lab is an integrated project. Students compare:
- exact dynamic programming,
- Q-learning,
- Dyna-Q,
- offline fitted Q iteration,
- conservative offline learning,
- behavior cloning.
The project asks students to produce a short report with:
- an MDP formulation,
- update equations,
- experimental settings,
- policy visualizations,
- learning curves,
- coverage diagnostics,
- a comparison table,
- an explanation of the results.
This final project is designed to show that reinforcement learning is not just a list of algorithms. It is a mathematical framework for learning, planning, and decision-making under uncertainty.
Closing perspective
Reinforcement learning is powerful because it combines three ideas:
- Prediction: estimate long-term consequences.
- Control: choose actions to improve outcomes.
- Learning: use data to improve predictions and decisions.
The mathematics of reinforcement learning gives us a language for these ideas.
This book develops that language from the beginning, connects it to computation, and shows how classical foundations lead naturally to modern reinforcement-learning methods.