Computer Labs
Google Colab-ready independent-study labs for The Mathematics of Reinforcement Learning
Computer Labs
This page summarizes the 28 computer labs for The Mathematics of Reinforcement Learning. Each lab is designed for independent study and can be opened directly in Google Colab.
Repository
The GitHub repository for the book is:
https://github.com/wanghemath/Book-MathRL
The lab notebooks are stored in:
labs/
The standard notebook naming pattern is:
labs/chapter-01-lab.ipynb
labs/chapter-02-lab.ipynb
...
labs/chapter-28-lab.ipynb
How to open a lab in Google Colab
Each lab can be opened by using a URL of the form:
https://colab.research.google.com/github/wanghemath/Book-MathRL/blob/main/labs/chapter-01-lab.ipynb
For example, Lab 1 can be opened here:
Purpose of the labs
The labs are designed to help students learn reinforcement learning by combining:
- mathematical explanation,
- small computational examples,
- readable Python code,
- visualizations,
- algorithm comparisons,
- interpretation questions,
- exercises,
- mini-projects.
Each lab is meant to be understandable even when studied independently. Background ideas are introduced before programming tasks.
Recommended workflow
For each lab:
- Read the mathematical background.
- Run all code cells once without modification.
- Identify the main update equation or optimization objective.
- Change one parameter and rerun the experiment.
- Explain how the output changed.
- Complete the exercises.
- Write a short paragraph connecting the code to the theory.
A good lab reflection should answer:
- What is being learned?
- What data does the algorithm use?
- What objective or fixed point is being approximated?
- What assumptions are being made?
- What can go wrong?
- How do the plots or tables show success or failure?
Full lab roadmap
| Lab | Module | Notebook | Colab | Main focus |
|---|---|---|---|---|
| 1 | Foundations | Introduction to Reinforcement Learning | Open in Colab | RL loop, trajectories, rewards, returns, policies, and first simulations. |
| 2 | Foundations | Markov Decision Processes | Open in Colab | Finite MDPs, transition probabilities, rewards, policies, and value functions. |
| 3 | Foundations | Bellman Equations and Dynamic Programming | Open in Colab | Bellman expectation/optimality equations, value iteration, and policy iteration. |
| 4 | Sample-based learning | Monte Carlo Methods | Open in Colab | Episode-based prediction and control using sampled returns. |
| 5 | Sample-based learning | Temporal-Difference Learning | Open in Colab | TD prediction, bootstrapping, TD errors, SARSA, and Q-learning preview. |
| 6 | Sample-based learning | SARSA and Q-Learning in Depth | Open in Colab | On-policy and off-policy TD control, risk, cliff walking, and exploration effects. |
| 7 | Exploration and bandits | Exploration and Exploitation | Open in Colab | Epsilon-greedy, decaying exploration, optimistic initialization, softmax, and UCB. |
| 8 | Exploration and bandits | Multi-Armed Bandits | Open in Colab | Regret, explore-then-commit, UCB, Thompson sampling, and bandit comparison. |
| 9 | Mathematical foundations | Stochastic Approximation | Open in Colab | Robbins–Monro updates, step-size schedules, noisy roots, and TD as stochastic approximation. |
| 10 | Approximation | Function Approximation | Open in Colab | Feature maps, regression viewpoint, linear approximation, RBF features, and semi-gradient TD. |
| 11 | Approximation | Linear Value Function Approximation | Open in Colab | Projection, normal equations, gradient Monte Carlo, semi-gradient TD, and LSTD ideas. |
| 12 | Policy optimization | Policy Gradient Methods | Open in Colab | Score-function identity, softmax policies, REINFORCE updates, and baselines. |
| 13 | Policy optimization | REINFORCE | Open in Colab | Full-return REINFORCE, reward-to-go, advantage-style updates, and variance reduction. |
| 14 | Policy optimization | Actor-Critic Methods | Open in Colab | Actor, critic, TD error as advantage, and one-step actor-critic learning. |
| 15 | Policy optimization | Natural Policy Gradient | Open in Colab | Fisher information, KL geometry, natural gradients, damping, and contextual bandits. |
| 16 | Policy optimization | Entropy-Regularized RL | Open in Colab | Entropy bonuses, temperature, soft Bellman equations, and soft value iteration. |
| 17 | Approximate/deep RL | Approximate Dynamic Programming | Open in Colab | Fitted value iteration, approximate planning, rollout improvement, and feature sensitivity. |
| 18 | Approximate/deep RL | Deep Q-Networks | Open in Colab | Neural Q-functions, replay buffers, target networks, and DQN-style learning. |
| 19 | Approximate/deep RL | Deep Policy Gradient Methods | Open in Colab | Neural stochastic policies, learned value baselines, and deep REINFORCE-style training. |
| 20 | Approximate/deep RL | Trust Region and PPO-Style Methods | Open in Colab | Policy ratios, clipped surrogate objectives, KL diagnostics, and PPO-style updates. |
| 21 | Approximate/deep RL | Continuous Control | Open in Colab | Continuous states/actions, Gaussian policies, policy gradients, and stabilizing controllers. |
| 22 | Model-based RL | Model-Based Reinforcement Learning | Open in Colab | Transition/reward model estimation, planning in learned models, and model-error diagnostics. |
| 23 | Model-based RL | Planning and Learning | Open in Colab | Dyna-Q, planning steps, prioritized sweeping, changing environments, and Dyna-Q+. |
| 24 | Offline/imitation/alignment | Offline Reinforcement Learning | Open in Colab | Fixed datasets, behavior policies, coverage, FQE, offline FQI, and conservative penalties. |
| 25 | Offline/imitation/alignment | Imitation Learning | Open in Colab | Expert demonstrations, behavior cloning, softmax classifiers, distribution shift, and DAgger. |
| 26 | Multi-agent/alignment | Multi-Agent Reinforcement Learning | Open in Colab | Matrix games, Markov games, independent Q-learning, coordination, and centralized training. |
| 27 | Multi-agent/alignment | Reinforcement Learning from Human Feedback | Open in Colab | Preference data, Bradley–Terry reward modeling, KL regularization, reward hacking, and DPO-style learning. |
| 28 | Capstone | Final Integrated Reinforcement Learning Project | Open in Colab | Capstone comparison of dynamic programming, Q-learning, Dyna-Q, offline RL, and imitation learning. |
Labs by module
Foundations
- Lab 1: Introduction to Reinforcement Learning — RL loop, trajectories, rewards, returns, policies, and first simulations. Open in Colab.
- Lab 2: Markov Decision Processes — Finite MDPs, transition probabilities, rewards, policies, and value functions. Open in Colab.
- Lab 3: Bellman Equations and Dynamic Programming — Bellman expectation/optimality equations, value iteration, and policy iteration. Open in Colab.
Sample-based learning
- Lab 4: Monte Carlo Methods — Episode-based prediction and control using sampled returns. Open in Colab.
- Lab 5: Temporal-Difference Learning — TD prediction, bootstrapping, TD errors, SARSA, and Q-learning preview. Open in Colab.
- Lab 6: SARSA and Q-Learning in Depth — On-policy and off-policy TD control, risk, cliff walking, and exploration effects. Open in Colab.
Exploration and bandits
- Lab 7: Exploration and Exploitation — Epsilon-greedy, decaying exploration, optimistic initialization, softmax, and UCB. Open in Colab.
- Lab 8: Multi-Armed Bandits — Regret, explore-then-commit, UCB, Thompson sampling, and bandit comparison. Open in Colab.
Mathematical foundations
- Lab 9: Stochastic Approximation — Robbins–Monro updates, step-size schedules, noisy roots, and TD as stochastic approximation. Open in Colab.
Approximation
- Lab 10: Function Approximation — Feature maps, regression viewpoint, linear approximation, RBF features, and semi-gradient TD. Open in Colab.
- Lab 11: Linear Value Function Approximation — Projection, normal equations, gradient Monte Carlo, semi-gradient TD, and LSTD ideas. Open in Colab.
Policy optimization
- Lab 12: Policy Gradient Methods — Score-function identity, softmax policies, REINFORCE updates, and baselines. Open in Colab.
- Lab 13: REINFORCE — Full-return REINFORCE, reward-to-go, advantage-style updates, and variance reduction. Open in Colab.
- Lab 14: Actor-Critic Methods — Actor, critic, TD error as advantage, and one-step actor-critic learning. Open in Colab.
- Lab 15: Natural Policy Gradient — Fisher information, KL geometry, natural gradients, damping, and contextual bandits. Open in Colab.
- Lab 16: Entropy-Regularized RL — Entropy bonuses, temperature, soft Bellman equations, and soft value iteration. Open in Colab.
Approximate/deep RL
- Lab 17: Approximate Dynamic Programming — Fitted value iteration, approximate planning, rollout improvement, and feature sensitivity. Open in Colab.
- Lab 18: Deep Q-Networks — Neural Q-functions, replay buffers, target networks, and DQN-style learning. Open in Colab.
- Lab 19: Deep Policy Gradient Methods — Neural stochastic policies, learned value baselines, and deep REINFORCE-style training. Open in Colab.
- Lab 20: Trust Region and PPO-Style Methods — Policy ratios, clipped surrogate objectives, KL diagnostics, and PPO-style updates. Open in Colab.
- Lab 21: Continuous Control — Continuous states/actions, Gaussian policies, policy gradients, and stabilizing controllers. Open in Colab.
Model-based RL
- Lab 22: Model-Based Reinforcement Learning — Transition/reward model estimation, planning in learned models, and model-error diagnostics. Open in Colab.
- Lab 23: Planning and Learning — Dyna-Q, planning steps, prioritized sweeping, changing environments, and Dyna-Q+. Open in Colab.
Offline/imitation/alignment
- Lab 24: Offline Reinforcement Learning — Fixed datasets, behavior policies, coverage, FQE, offline FQI, and conservative penalties. Open in Colab.
- Lab 25: Imitation Learning — Expert demonstrations, behavior cloning, softmax classifiers, distribution shift, and DAgger. Open in Colab.
Multi-agent/alignment
- Lab 26: Multi-Agent Reinforcement Learning — Matrix games, Markov games, independent Q-learning, coordination, and centralized training. Open in Colab.
- Lab 27: Reinforcement Learning from Human Feedback — Preference data, Bradley–Terry reward modeling, KL regularization, reward hacking, and DPO-style learning. Open in Colab.
Capstone
- Lab 28: Final Integrated Reinforcement Learning Project — Capstone comparison of dynamic programming, Q-learning, Dyna-Q, offline RL, and imitation learning. Open in Colab.
Lab sequence
The labs are intentionally cumulative.
Labs 1–3: MDP and Bellman foundations
These labs establish the basic mathematical language:
\[ S_t,\quad A_t,\quad R_{t+1},\quad S_{t+1}, \]
policies, value functions, transition probabilities, and Bellman equations.
Labs 4–6: Learning from sampled experience
These labs move from exact dynamic programming to sample-based learning:
\[ V(S_t) \leftarrow V(S_t) + \alpha [ R_{t+1}+\gamma V(S_{t+1})-V(S_t) ]. \]
Students compare Monte Carlo, TD learning, SARSA, and Q-learning.
Labs 7–9: Exploration, bandits, and stochastic approximation
These labs study how agents gather information and how noisy recursive updates converge.
Topics include regret, UCB, Thompson sampling, and Robbins–Monro step-size conditions.
Labs 10–11: Function approximation
These labs introduce feature-based value functions:
\[ V(s;w)=\phi(s)^T w. \]
Students study approximation error, scaling, projection, gradient Monte Carlo, semi-gradient TD, and LSTD-style ideas.
Labs 12–16: Policy optimization
These labs study direct policy optimization:
\[ \nabla_\theta J(\theta) = E[ G_t\nabla_\theta \log \pi_\theta(A_t\mid S_t) ]. \]
Students implement REINFORCE, actor-critic methods, natural policy gradient, and entropy-regularized RL.
Labs 17–21: Approximate and deep reinforcement learning
These labs introduce approximate dynamic programming, DQN-style ideas, deep policy gradients, PPO-style clipping, and continuous control.
The goal is to connect classical RL equations to modern function-approximation methods.
Labs 22–23: Model-based RL and planning
These labs study model estimation and planning.
Students estimate:
\[ \widehat P(s'\mid s,a), \qquad \widehat R(s,a), \]
then plan using value iteration or Dyna-Q-style simulated updates.
Labs 24–27: Modern advanced topics
These labs introduce offline RL, imitation learning, multi-agent RL, and RLHF.
The main themes are:
- data coverage,
- distribution shift,
- preference learning,
- coordination,
- conservative learning,
- alignment-style regularization.
Lab 28: Final integrated project
The final lab is a capstone project. Students compare multiple RL approaches on one environment:
- exact dynamic programming,
- Q-learning,
- Dyna-Q,
- offline fitted Q iteration,
- conservative offline learning,
- behavior cloning.
The final deliverable is a short report with equations, experiments, plots, diagnostics, and conclusions.
Suggested pacing
A semester course could use the labs as follows:
| Week range | Labs | Theme |
|---|---|---|
| Weeks 1–2 | 1–3 | MDPs, Bellman equations, dynamic programming |
| Weeks 3–4 | 4–6 | Monte Carlo and TD control |
| Weeks 5–6 | 7–9 | Exploration, bandits, stochastic approximation |
| Weeks 7–8 | 10–13 | Function approximation and policy gradients |
| Weeks 9–10 | 14–18 | Actor-critic, entropy, ADP, and DQN |
| Weeks 11–12 | 19–23 | Deep policy gradients, PPO-style methods, continuous control, model-based RL |
| Weeks 13–14 | 24–27 | Offline RL, imitation learning, multi-agent RL, RLHF |
| Final project | 28 | Integrated capstone |
For a shorter course, instructors may select:
- Labs 1–6 for a classical RL introduction,
- Labs 7–9 for exploration and bandits,
- Labs 10–16 for policy optimization,
- Labs 18–23 for deep and model-based RL,
- Labs 24–28 for modern applied topics and projects.
Lab report template
For mini-projects and final projects, students may use the following structure.
1. Problem
Describe the environment or decision problem:
- state space,
- action space,
- reward function,
- transition uncertainty,
- terminal states.
2. Method
State the algorithm and its update equation.
For example, Q-learning uses:
\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]
3. Experiment
Report:
- number of episodes,
- learning rate,
- discount factor,
- exploration schedule,
- random seed,
- evaluation method.
4. Results
Include:
- a learning curve,
- a policy or value visualization,
- a summary table,
- a short interpretation.
5. Diagnosis
Explain success or failure using concepts such as:
- exploration,
- variance,
- bias,
- coverage,
- model error,
- bootstrapping,
- distribution shift,
- reward misspecification,
- nonstationarity.
6. Conclusion
Summarize the main lesson from the lab.
Common implementation conventions
Most labs use:
import numpy as np
import pandas as pd
import matplotlib.pyplot as pltMany labs define small environments directly in the notebook. This avoids heavy dependencies and keeps the mathematical structure visible.
Common variables include:
| Variable | Meaning |
|---|---|
env |
environment object |
s |
current state |
a |
action |
r |
reward |
sp |
next state |
done |
episode termination flag |
gamma |
discount factor |
alpha |
learning rate |
epsilon |
exploration probability |
Q |
action-value table or approximation |
V |
value function |
policy |
deterministic or stochastic policy |
returns |
episode returns |
history |
dictionary of learning diagnostics |
Assessment ideas
The labs can be used for:
- homework assignments,
- in-class demonstrations,
- project preparation,
- independent study,
- final project scaffolding.
Possible grading dimensions:
| Category | Evidence |
|---|---|
| Mathematical understanding | Correct equations and explanations |
| Code execution | Notebook runs without errors |
| Experimentation | Student changes parameters and compares outcomes |
| Visualization | Clear plots and tables |
| Interpretation | Student explains why results changed |
| Extension | Student adds a meaningful modification |
| Communication | Final write-up is clear and concise |
Final note
The purpose of these labs is not only to teach Python implementations of reinforcement-learning algorithms. The deeper goal is to help students see the unity between:
- Bellman equations,
- stochastic approximation,
- policy optimization,
- exploration,
- planning,
- function approximation,
- offline data,
- demonstrations,
- preferences,
- and multi-agent interaction.
Together, the labs turn reinforcement learning into a mathematical and computational language for sequential decision-making under uncertainty.