Computer Labs

Google Colab-ready independent-study labs for The Mathematics of Reinforcement Learning

Computer Labs

Note

This page summarizes the 28 computer labs for The Mathematics of Reinforcement Learning. Each lab is designed for independent study and can be opened directly in Google Colab.

Repository

The GitHub repository for the book is:

https://github.com/wanghemath/Book-MathRL

The lab notebooks are stored in:

labs/

The standard notebook naming pattern is:

labs/chapter-01-lab.ipynb
labs/chapter-02-lab.ipynb
...
labs/chapter-28-lab.ipynb

How to open a lab in Google Colab

Each lab can be opened by using a URL of the form:

https://colab.research.google.com/github/wanghemath/Book-MathRL/blob/main/labs/chapter-01-lab.ipynb

For example, Lab 1 can be opened here:

Open Lab 1 in Google Colab

Purpose of the labs

The labs are designed to help students learn reinforcement learning by combining:

  • mathematical explanation,
  • small computational examples,
  • readable Python code,
  • visualizations,
  • algorithm comparisons,
  • interpretation questions,
  • exercises,
  • mini-projects.

Each lab is meant to be understandable even when studied independently. Background ideas are introduced before programming tasks.

Full lab roadmap

Lab Module Notebook Colab Main focus
1 Foundations Introduction to Reinforcement Learning Open in Colab RL loop, trajectories, rewards, returns, policies, and first simulations.
2 Foundations Markov Decision Processes Open in Colab Finite MDPs, transition probabilities, rewards, policies, and value functions.
3 Foundations Bellman Equations and Dynamic Programming Open in Colab Bellman expectation/optimality equations, value iteration, and policy iteration.
4 Sample-based learning Monte Carlo Methods Open in Colab Episode-based prediction and control using sampled returns.
5 Sample-based learning Temporal-Difference Learning Open in Colab TD prediction, bootstrapping, TD errors, SARSA, and Q-learning preview.
6 Sample-based learning SARSA and Q-Learning in Depth Open in Colab On-policy and off-policy TD control, risk, cliff walking, and exploration effects.
7 Exploration and bandits Exploration and Exploitation Open in Colab Epsilon-greedy, decaying exploration, optimistic initialization, softmax, and UCB.
8 Exploration and bandits Multi-Armed Bandits Open in Colab Regret, explore-then-commit, UCB, Thompson sampling, and bandit comparison.
9 Mathematical foundations Stochastic Approximation Open in Colab Robbins–Monro updates, step-size schedules, noisy roots, and TD as stochastic approximation.
10 Approximation Function Approximation Open in Colab Feature maps, regression viewpoint, linear approximation, RBF features, and semi-gradient TD.
11 Approximation Linear Value Function Approximation Open in Colab Projection, normal equations, gradient Monte Carlo, semi-gradient TD, and LSTD ideas.
12 Policy optimization Policy Gradient Methods Open in Colab Score-function identity, softmax policies, REINFORCE updates, and baselines.
13 Policy optimization REINFORCE Open in Colab Full-return REINFORCE, reward-to-go, advantage-style updates, and variance reduction.
14 Policy optimization Actor-Critic Methods Open in Colab Actor, critic, TD error as advantage, and one-step actor-critic learning.
15 Policy optimization Natural Policy Gradient Open in Colab Fisher information, KL geometry, natural gradients, damping, and contextual bandits.
16 Policy optimization Entropy-Regularized RL Open in Colab Entropy bonuses, temperature, soft Bellman equations, and soft value iteration.
17 Approximate/deep RL Approximate Dynamic Programming Open in Colab Fitted value iteration, approximate planning, rollout improvement, and feature sensitivity.
18 Approximate/deep RL Deep Q-Networks Open in Colab Neural Q-functions, replay buffers, target networks, and DQN-style learning.
19 Approximate/deep RL Deep Policy Gradient Methods Open in Colab Neural stochastic policies, learned value baselines, and deep REINFORCE-style training.
20 Approximate/deep RL Trust Region and PPO-Style Methods Open in Colab Policy ratios, clipped surrogate objectives, KL diagnostics, and PPO-style updates.
21 Approximate/deep RL Continuous Control Open in Colab Continuous states/actions, Gaussian policies, policy gradients, and stabilizing controllers.
22 Model-based RL Model-Based Reinforcement Learning Open in Colab Transition/reward model estimation, planning in learned models, and model-error diagnostics.
23 Model-based RL Planning and Learning Open in Colab Dyna-Q, planning steps, prioritized sweeping, changing environments, and Dyna-Q+.
24 Offline/imitation/alignment Offline Reinforcement Learning Open in Colab Fixed datasets, behavior policies, coverage, FQE, offline FQI, and conservative penalties.
25 Offline/imitation/alignment Imitation Learning Open in Colab Expert demonstrations, behavior cloning, softmax classifiers, distribution shift, and DAgger.
26 Multi-agent/alignment Multi-Agent Reinforcement Learning Open in Colab Matrix games, Markov games, independent Q-learning, coordination, and centralized training.
27 Multi-agent/alignment Reinforcement Learning from Human Feedback Open in Colab Preference data, Bradley–Terry reward modeling, KL regularization, reward hacking, and DPO-style learning.
28 Capstone Final Integrated Reinforcement Learning Project Open in Colab Capstone comparison of dynamic programming, Q-learning, Dyna-Q, offline RL, and imitation learning.

Labs by module

Foundations

Sample-based learning

Exploration and bandits

Mathematical foundations

Approximation

Policy optimization

Approximate/deep RL

Model-based RL

Offline/imitation/alignment

Multi-agent/alignment

Capstone

Lab sequence

The labs are intentionally cumulative.

Labs 1–3: MDP and Bellman foundations

These labs establish the basic mathematical language:

\[ S_t,\quad A_t,\quad R_{t+1},\quad S_{t+1}, \]

policies, value functions, transition probabilities, and Bellman equations.

Labs 4–6: Learning from sampled experience

These labs move from exact dynamic programming to sample-based learning:

\[ V(S_t) \leftarrow V(S_t) + \alpha [ R_{t+1}+\gamma V(S_{t+1})-V(S_t) ]. \]

Students compare Monte Carlo, TD learning, SARSA, and Q-learning.

Labs 7–9: Exploration, bandits, and stochastic approximation

These labs study how agents gather information and how noisy recursive updates converge.

Topics include regret, UCB, Thompson sampling, and Robbins–Monro step-size conditions.

Labs 10–11: Function approximation

These labs introduce feature-based value functions:

\[ V(s;w)=\phi(s)^T w. \]

Students study approximation error, scaling, projection, gradient Monte Carlo, semi-gradient TD, and LSTD-style ideas.

Labs 12–16: Policy optimization

These labs study direct policy optimization:

\[ \nabla_\theta J(\theta) = E[ G_t\nabla_\theta \log \pi_\theta(A_t\mid S_t) ]. \]

Students implement REINFORCE, actor-critic methods, natural policy gradient, and entropy-regularized RL.

Labs 17–21: Approximate and deep reinforcement learning

These labs introduce approximate dynamic programming, DQN-style ideas, deep policy gradients, PPO-style clipping, and continuous control.

The goal is to connect classical RL equations to modern function-approximation methods.

Labs 22–23: Model-based RL and planning

These labs study model estimation and planning.

Students estimate:

\[ \widehat P(s'\mid s,a), \qquad \widehat R(s,a), \]

then plan using value iteration or Dyna-Q-style simulated updates.

Labs 24–27: Modern advanced topics

These labs introduce offline RL, imitation learning, multi-agent RL, and RLHF.

The main themes are:

  • data coverage,
  • distribution shift,
  • preference learning,
  • coordination,
  • conservative learning,
  • alignment-style regularization.

Lab 28: Final integrated project

The final lab is a capstone project. Students compare multiple RL approaches on one environment:

  • exact dynamic programming,
  • Q-learning,
  • Dyna-Q,
  • offline fitted Q iteration,
  • conservative offline learning,
  • behavior cloning.

The final deliverable is a short report with equations, experiments, plots, diagnostics, and conclusions.

Suggested pacing

A semester course could use the labs as follows:

Week range Labs Theme
Weeks 1–2 1–3 MDPs, Bellman equations, dynamic programming
Weeks 3–4 4–6 Monte Carlo and TD control
Weeks 5–6 7–9 Exploration, bandits, stochastic approximation
Weeks 7–8 10–13 Function approximation and policy gradients
Weeks 9–10 14–18 Actor-critic, entropy, ADP, and DQN
Weeks 11–12 19–23 Deep policy gradients, PPO-style methods, continuous control, model-based RL
Weeks 13–14 24–27 Offline RL, imitation learning, multi-agent RL, RLHF
Final project 28 Integrated capstone

For a shorter course, instructors may select:

  • Labs 1–6 for a classical RL introduction,
  • Labs 7–9 for exploration and bandits,
  • Labs 10–16 for policy optimization,
  • Labs 18–23 for deep and model-based RL,
  • Labs 24–28 for modern applied topics and projects.

Lab report template

For mini-projects and final projects, students may use the following structure.

1. Problem

Describe the environment or decision problem:

  • state space,
  • action space,
  • reward function,
  • transition uncertainty,
  • terminal states.

2. Method

State the algorithm and its update equation.

For example, Q-learning uses:

\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]

3. Experiment

Report:

  • number of episodes,
  • learning rate,
  • discount factor,
  • exploration schedule,
  • random seed,
  • evaluation method.

4. Results

Include:

  • a learning curve,
  • a policy or value visualization,
  • a summary table,
  • a short interpretation.

5. Diagnosis

Explain success or failure using concepts such as:

  • exploration,
  • variance,
  • bias,
  • coverage,
  • model error,
  • bootstrapping,
  • distribution shift,
  • reward misspecification,
  • nonstationarity.

6. Conclusion

Summarize the main lesson from the lab.

Common implementation conventions

Most labs use:

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

Many labs define small environments directly in the notebook. This avoids heavy dependencies and keeps the mathematical structure visible.

Common variables include:

Variable Meaning
env environment object
s current state
a action
r reward
sp next state
done episode termination flag
gamma discount factor
alpha learning rate
epsilon exploration probability
Q action-value table or approximation
V value function
policy deterministic or stochastic policy
returns episode returns
history dictionary of learning diagnostics

Assessment ideas

The labs can be used for:

  • homework assignments,
  • in-class demonstrations,
  • project preparation,
  • independent study,
  • final project scaffolding.

Possible grading dimensions:

Category Evidence
Mathematical understanding Correct equations and explanations
Code execution Notebook runs without errors
Experimentation Student changes parameters and compares outcomes
Visualization Clear plots and tables
Interpretation Student explains why results changed
Extension Student adds a meaningful modification
Communication Final write-up is clear and concise

Final note

The purpose of these labs is not only to teach Python implementations of reinforcement-learning algorithms. The deeper goal is to help students see the unity between:

  • Bellman equations,
  • stochastic approximation,
  • policy optimization,
  • exploration,
  • planning,
  • function approximation,
  • offline data,
  • demonstrations,
  • preferences,
  • and multi-agent interaction.

Together, the labs turn reinforcement learning into a mathematical and computational language for sequential decision-making under uncertainty.