Project Guide

Final projects and mini-projects for The Mathematics of Reinforcement Learning

Project Guide

Note

This guide explains how to design, complete, and write up reinforcement-learning projects for The Mathematics of Reinforcement Learning. It can be used for mini-projects, lab extensions, course projects, and the final integrated project.

Purpose

A reinforcement-learning project should do more than run an algorithm. A strong project should connect:

  • a mathematical formulation,
  • an algorithmic idea,
  • a computational implementation,
  • an experimental comparison,
  • a diagnosis of success or failure,
  • a clear written conclusion.

The goal is to practice reinforcement learning as a mathematical and computational discipline.

A good project answers questions such as:

  • What is the sequential decision problem?
  • What is the state space?
  • What are the actions?
  • What reward is being optimized?
  • What data or interaction is available?
  • What algorithm is used?
  • What assumptions does the algorithm make?
  • How is performance evaluated?
  • Why did the method succeed or fail?

Repository

The GitHub repository for the book is:

https://github.com/wanghemath/Book-MathRL

The labs are stored in:

labs/

The final integrated lab is:

labs/chapter-28-lab.ipynb

It can be opened in Google Colab:

https://colab.research.google.com/github/wanghemath/Book-MathRL/blob/main/labs/chapter-28-lab.ipynb

What counts as a project?

A project may be one of the following:

  1. Algorithm comparison project
    Compare several reinforcement-learning algorithms on one environment.

  2. Environment-design project
    Modify or create an environment and study how different algorithms behave.

  3. Theory-to-code project
    Start from an equation, implement the algorithm, and analyze convergence or failure.

  4. Modern RL topic project
    Study offline RL, imitation learning, multi-agent RL, model-based RL, or RLHF.

  5. Applied modeling project
    Formulate an applied problem as an MDP or contextual decision problem and solve a simplified version.

  6. Final integrated project
    Use Lab 28 as the base and extend it in a meaningful direction.

Core project structure

Every project should include the following parts.

1. Problem formulation

Describe the sequential decision problem.

Include:

  • state space \(S\),
  • action space \(A\),
  • reward function \(R\),
  • transition structure,
  • discount factor \(\gamma\),
  • terminal states, if any,
  • policy class.

A finite MDP can be summarized as:

\[ (S,A,P,R,\gamma). \]

A policy is a rule:

\[ \pi(a\mid s). \]

The value of a policy is:

\[ V^\pi(s) = E_\pi \left[ \sum_{t=0}^{\infty} \gamma^t R_{t+1} \mid S_0=s \right]. \]

2. Algorithm description

State the method clearly.

For example, Q-learning uses:

\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]

SARSA uses:

\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma Q(s',a')-Q(s,a) \right]. \]

REINFORCE uses:

\[ \theta \leftarrow \theta + \alpha G_t \nabla_\theta \log \pi_\theta(A_t\mid S_t). \]

Dyna-Q combines real updates with simulated model-based updates.

Offline fitted Q iteration uses a fixed dataset:

\[ D=\{(s_i,a_i,r_i,s_i',d_i)\}_{i=1}^N. \]

3. Experimental design

Report all important experimental settings:

  • number of episodes,
  • number of random seeds,
  • learning rate,
  • discount factor,
  • exploration schedule,
  • function approximation features,
  • neural-network size, if used,
  • number of planning steps, if model-based,
  • offline dataset size, if offline,
  • number of expert demonstrations, if imitation learning,
  • evaluation episodes.

A project should be reproducible. If another student reads your report, they should be able to rerun the experiment.

4. Evaluation metrics

Choose metrics appropriate for the task.

Common metrics include:

Metric Meaning
Mean return Average discounted or undiscounted return
Success rate Fraction of episodes reaching the goal
Failure/trap rate Fraction of episodes reaching bad terminal states
Episode length Number of steps before termination
Regret Loss relative to the best action or policy
Bellman error Size of Bellman residual or TD error
Policy entropy Randomness of a stochastic policy
KL divergence Distance from a reference policy
Coverage Number of visited state-action pairs
Model error Difference between true and learned transition models
Preference accuracy Accuracy of a reward model on comparisons

5. Visualization

Include at least two meaningful visualizations.

Good options include:

  • learning curve,
  • rolling average return,
  • success-rate curve,
  • policy arrow plot,
  • value-function heatmap,
  • state-visitation heatmap,
  • regret curve,
  • action-probability plot,
  • reward-model scatter plot,
  • offline coverage plot,
  • comparison bar chart.

A plot should not appear alone. Explain what it shows.

6. Diagnosis

A good project explains why the results happened.

Possible diagnostic themes:

  • insufficient exploration,
  • large variance,
  • biased value estimates,
  • bootstrapping instability,
  • poor step-size choice,
  • bad features,
  • distribution shift,
  • poor offline coverage,
  • model error,
  • reward misspecification,
  • nonstationarity from other agents,
  • reward hacking,
  • excessive or insufficient regularization.

7. Conclusion

The conclusion should answer:

  • Which method worked best?
  • Why did it work best?
  • What failed?
  • What would you try next?
  • What did the experiment teach about reinforcement learning?

Project proposal template

Before starting a final project, write a short proposal.

Title

Give the project a clear title.

Example:

Comparing Q-Learning, Dyna-Q, and Offline FQI in a Stochastic Grid-World

Question

State the main research question.

Examples:

  • Does planning improve sample efficiency in this environment?
  • How does offline dataset coverage affect learned policy quality?
  • Does entropy regularization improve exploration?
  • How sensitive is Q-learning to the learning rate?
  • Can behavior cloning match expert performance with limited demonstrations?

Environment

Describe the environment:

  • state space,
  • action space,
  • reward,
  • transition randomness,
  • terminal states.

Methods

List the methods you will compare.

Examples:

  • value iteration,
  • Q-learning,
  • SARSA,
  • Dyna-Q,
  • REINFORCE,
  • actor-critic,
  • offline fitted Q iteration,
  • behavior cloning.

Metrics

List the metrics you will report.

Examples:

  • mean return,
  • success rate,
  • trap rate,
  • episode length,
  • coverage,
  • policy disagreement,
  • regret.

Expected result

State your hypothesis.

Example:

I expect Dyna-Q to learn faster than Q-learning because it uses planning updates from a learned model.

Extension

Describe what is new beyond the base lab.

Examples:

  • changed environment,
  • new hyperparameter study,
  • new diagnostic plot,
  • new algorithm variant,
  • new comparison table,
  • robustness experiment.

Final report template

Use the following structure for the final write-up.

1. Abstract

Write a short summary of the project in 5–8 sentences.

Include:

  • the problem,
  • the algorithms,
  • the main result,
  • the main conclusion.

2. Introduction

Explain the motivation.

Example:

This project studies how model-based planning changes sample efficiency in a stochastic navigation task.

3. MDP formulation

Define the MDP:

\[ (S,A,P,R,\gamma). \]

Include:

  • states,
  • actions,
  • rewards,
  • transitions,
  • terminal states,
  • discount factor.

4. Algorithms

Describe each algorithm.

For each method, include:

  • the main update equation,
  • the data used,
  • the policy used for action selection,
  • any important hyperparameters.

5. Experimental setup

Report:

  • number of training episodes,
  • number of evaluation episodes,
  • random seeds,
  • learning rates,
  • exploration schedule,
  • planning steps,
  • offline dataset size,
  • demonstration size.

6. Results

Include:

  • a comparison table,
  • at least one learning curve,
  • at least one policy or value visualization,
  • at least one diagnostic plot.

7. Discussion

Explain the results.

Address:

  • why some methods worked better,
  • why some methods failed,
  • whether the results match your hypothesis,
  • what limitations remain.

8. Conclusion

Summarize the main lesson.

9. Appendix

Optional. Include extra plots, code snippets, or additional experiments.

Suggested project topics

Topic 1: Q-learning versus SARSA

Question

How do on-policy and off-policy TD control methods differ in risky environments?

Base labs

  • Lab 5: Temporal-Difference Learning
  • Lab 6: SARSA and Q-Learning in Depth

Suggested experiment

Use a cliff-walking or trap-navigation environment.

Compare:

  • SARSA,
  • Q-learning,
  • Expected SARSA.

Report:

  • learning curves,
  • final policies,
  • average return,
  • failure rate.

Key concept

SARSA accounts for exploratory behavior, while Q-learning learns a greedy target policy.

Topic 2: Exploration strategies

Question

Which exploration strategy works best in a sparse-reward environment?

Base labs

  • Lab 7: Exploration and Exploitation
  • Lab 8: Multi-Armed Bandits

Suggested experiment

Compare:

  • fixed epsilon-greedy,
  • decaying epsilon-greedy,
  • optimistic initialization,
  • softmax exploration,
  • UCB.

Report:

  • cumulative reward,
  • regret,
  • optimal-action frequency,
  • sensitivity to parameters.

Key concept

Exploration is a statistical problem and a control problem at the same time.

Topic 3: Step-size schedules and stochastic approximation

Question

How does the learning-rate schedule affect convergence?

Base labs

  • Lab 9: Stochastic Approximation

Suggested experiment

Compare:

\[ \alpha_t = 0.1, \qquad \alpha_t = \frac{1}{t}, \qquad \alpha_t = \frac{1}{\sqrt{t}}, \qquad \alpha_t = \frac{0.5}{10+t}. \]

Report:

  • convergence curves,
  • variance across seeds,
  • final error.

Key concept

Step sizes control the tradeoff between adaptation and convergence.

Topic 4: Function approximation

Question

How do features affect value-function approximation?

Base labs

  • Lab 10: Function Approximation
  • Lab 11: Linear Value Function Approximation

Suggested experiment

Compare:

  • state aggregation,
  • polynomial features,
  • radial basis functions,
  • Fourier features.

Report:

  • approximation error,
  • learning curves,
  • final policy quality,
  • feature sensitivity.

Key concept

Function approximation introduces generalization and approximation error.

Topic 5: Policy-gradient variance reduction

Question

How much do reward-to-go and baselines reduce policy-gradient variance?

Base labs

  • Lab 12: Policy Gradient Methods
  • Lab 13: REINFORCE
  • Lab 14: Actor-Critic Methods

Suggested experiment

Compare:

  • full-return REINFORCE,
  • reward-to-go REINFORCE,
  • REINFORCE with baseline,
  • actor-critic.

Report:

  • learning curves,
  • variance across seeds,
  • final success rate,
  • policy entropy.

Key concept

Baselines reduce variance without changing the expected policy gradient.

Topic 6: Entropy regularization

Question

Does entropy regularization improve exploration and prevent premature policy collapse?

Base labs

  • Lab 16: Entropy-Regularized RL

Suggested experiment

Train policies with different entropy temperatures:

\[ \tau = 0,\quad 0.01,\quad 0.05,\quad 0.1,\quad 0.5. \]

Report:

  • return,
  • entropy,
  • action probabilities,
  • final policy.

Key concept

Entropy regularization encourages stochasticity and can improve exploration.

Topic 7: DQN-style stabilizers

Question

How do replay buffers and target networks stabilize neural Q-learning?

Base labs

  • Lab 18: Deep Q-Networks

Suggested experiment

Compare:

  • no replay buffer,
  • replay buffer,
  • target network,
  • replay buffer plus target network.

Report:

  • TD loss,
  • return,
  • final policy,
  • instability or divergence.

Key concept

Deep RL requires stabilization because bootstrapping and function approximation can interact badly.

Topic 8: PPO-style clipping

Question

How does clipping affect policy-gradient stability?

Base labs

  • Lab 20: Trust Region and PPO-Style Methods

Suggested experiment

Compare clipping parameters:

\[ \epsilon_{\text{clip}} = 0.05,\quad 0.1,\quad 0.2,\quad 0.4. \]

Report:

  • return,
  • KL divergence,
  • clipping fraction,
  • policy entropy.

Key concept

PPO-style clipping limits destructive policy updates.

Topic 9: Continuous control

Question

How does Gaussian exploration noise affect continuous control?

Base labs

  • Lab 21: Continuous Control

Suggested experiment

Compare Gaussian standard deviations:

\[ \sigma = 0.1,\quad 0.3,\quad 0.6,\quad 1.0. \]

Report:

  • return,
  • state magnitude,
  • action magnitude,
  • learned mean controller.

Key concept

Continuous control requires balancing exploration noise with stable control.

Topic 10: Model-based RL

Question

How much data is needed to learn a useful model?

Base labs

  • Lab 22: Model-Based Reinforcement Learning

Suggested experiment

Collect datasets of different sizes. Estimate transition models and plan inside them.

Report:

  • model error,
  • policy success rate,
  • state-action coverage,
  • learned policy plots.

Key concept

Planning is only as good as the learned model.

Topic 11: Dyna-Q and planning steps

Question

How many planning steps are useful?

Base labs

  • Lab 23: Planning and Learning

Suggested experiment

Compare:

\[ n_{\text{planning}} = 0,\quad 1,\quad 5,\quad 20,\quad 50. \]

Report:

  • learning curves,
  • sample efficiency,
  • computation cost,
  • final policy.

Key concept

Planning reuses experience but may waste computation or amplify model errors.

Topic 12: Offline RL and coverage

Question

How does offline dataset coverage affect learned policy quality?

Base labs

  • Lab 24: Offline Reinforcement Learning

Suggested experiment

Generate:

  • random dataset,
  • medium-quality dataset,
  • expert dataset.

Compare:

  • naive offline FQI,
  • supported-action FQI,
  • count-penalized FQI.

Report:

  • coverage,
  • learned policies,
  • return,
  • out-of-distribution action diagnostics.

Key concept

Offline RL is fundamentally limited by dataset support.

Topic 13: Imitation learning and DAgger

Question

How does DAgger reduce distribution shift in behavior cloning?

Base labs

  • Lab 25: Imitation Learning

Suggested experiment

Compare:

  • behavior cloning with few demonstrations,
  • behavior cloning with many demonstrations,
  • DAgger-style dataset aggregation.

Report:

  • state coverage,
  • success rate,
  • final policy,
  • learner-visited unseen states.

Key concept

DAgger labels states induced by the learner, not only states visited by the expert.

Topic 14: Multi-agent coordination

Question

How stable is independent Q-learning in multi-agent games?

Base labs

  • Lab 26: Multi-Agent Reinforcement Learning

Suggested experiment

Compare learning in:

  • coordination game,
  • matching pennies,
  • prisoner’s dilemma,
  • cooperative grid task.

Report:

  • action frequencies,
  • rewards,
  • equilibrium behavior,
  • seed sensitivity.

Key concept

Other learning agents make the environment nonstationary.

Topic 15: RLHF reward modeling

Question

How does preference-data quality affect reward-model optimization?

Base labs

  • Lab 27: Reinforcement Learning from Human Feedback

Suggested experiment

Compare:

  • small versus large preference datasets,
  • unbiased versus biased preference data,
  • weak versus strong KL regularization,
  • RLHF-style reward optimization versus DPO-style updates.

Report:

  • reward-model rank correlation,
  • KL versus reward plots,
  • action distributions,
  • reward hacking diagnostics.

Key concept

Optimizing a learned reward can fail when the reward model is biased or inaccurate.

Topic 16: Final integrated comparison

Question

Which family of methods works best in a stochastic navigation task?

Base lab

  • Lab 28: Final Integrated Reinforcement Learning Project

Suggested experiment

Compare:

  • exact dynamic programming,
  • Q-learning,
  • Dyna-Q,
  • offline fitted Q iteration,
  • conservative offline learning,
  • behavior cloning.

Report:

  • final policy plots,
  • learning curves,
  • success rates,
  • coverage diagnostics,
  • policy disagreement with the dynamic-programming benchmark.

Key concept

Different RL methods use different information: known models, interaction, offline data, demonstrations, or learned models.

Reproducibility checklist

Before submitting a project, check the following.

Suggested grading rubric

Category Excellent Satisfactory Needs improvement
MDP formulation State, action, reward, transition, and objective are clear Most elements are clear Formulation is incomplete
Mathematical explanation Correct equations and interpretation Some equations included Little connection to math
Implementation Correct, readable, reproducible Mostly correct Does not run or is unclear
Experimental design Fair comparison with clear settings Basic comparison Missing controls or settings
Evaluation Multiple metrics and meaningful plots Some metrics Weak or missing evaluation
Diagnosis Explains why results occurred Some interpretation Mostly descriptive
Extension Meaningful original modification Small modification No meaningful extension
Communication Clear report with strong organization Understandable report Difficult to follow

Common project mistakes

Avoid the following.

Mistake 1: Only reporting one run

RL algorithms are stochastic. Use multiple seeds when possible.

Mistake 2: Comparing methods unfairly

Use the same environment and evaluation function for all methods.

Mistake 3: Confusing training and evaluation

Exploration is often used during training. Evaluation should usually use the learned greedy or mean policy.

Mistake 4: Ignoring coverage

Offline RL and imitation learning depend heavily on what states and actions appear in the data.

Mistake 5: Showing plots without explanation

Every plot should have a short interpretation.

Mistake 6: Reporting success without diagnosis

A good project explains why the method worked.

Mistake 7: Changing too many things at once

Change one variable at a time when possible.

Mini-project format

For smaller assignments, use this shorter format.

Mini-project report

  1. Question: What are you testing?
  2. Method: What algorithm did you use?
  3. Equation: What is the key update or objective?
  4. Experiment: What parameter or environment feature did you change?
  5. Result: What plot or table summarizes the result?
  6. Interpretation: Why did this happen?
  7. Conclusion: What did you learn?

A mini-project can be 1–2 pages plus code.

Final project format

For the final project, use a longer format.

Recommended length:

  • 4–8 pages,
  • plus notebook,
  • plus figures and tables.

Required components:

  1. abstract,
  2. MDP formulation,
  3. algorithm descriptions,
  4. experimental setup,
  5. results,
  6. diagnosis,
  7. conclusion,
  8. reproducibility notes.

Final project presentation

A short presentation may use the following structure.

Slide 1: Title and question

State the project question.

Slide 2: Environment

Show the state/action/reward setup.

Slide 3: Algorithms

List the methods compared.

Slide 4: Main equations

Show the key update equations.

Slide 5: Results table

Compare final metrics.

Slide 6: Learning curves

Show training behavior.

Slide 7: Diagnostics

Show coverage, model error, policy disagreement, or reward-model issue.

Slide 8: Conclusion

State the main lesson.

A strong final conclusion

A strong conclusion should be specific.

Weak conclusion:

Dyna-Q worked well.

Stronger conclusion:

Dyna-Q reached a high success rate in fewer episodes than Q-learning because planning updates reused each real transition several times. However, its advantage depended on the learned model being accurate; in the changed-environment experiment, outdated model transitions slowed adaptation.

Closing advice

A reinforcement-learning project is successful when it teaches a clear lesson.

The best projects are not necessarily the most complicated. A simple environment with careful comparisons and thoughtful diagnosis is often better than a large experiment with unclear conclusions.

Focus on the mathematical story:

  • What is the objective?
  • What information does the algorithm use?
  • What approximation is made?
  • What evidence shows success?
  • What evidence explains failure?

That is the core of doing reinforcement learning well.