Project Guide
Final projects and mini-projects for The Mathematics of Reinforcement Learning
Project Guide
This guide explains how to design, complete, and write up reinforcement-learning projects for The Mathematics of Reinforcement Learning. It can be used for mini-projects, lab extensions, course projects, and the final integrated project.
Purpose
A reinforcement-learning project should do more than run an algorithm. A strong project should connect:
- a mathematical formulation,
- an algorithmic idea,
- a computational implementation,
- an experimental comparison,
- a diagnosis of success or failure,
- a clear written conclusion.
The goal is to practice reinforcement learning as a mathematical and computational discipline.
A good project answers questions such as:
- What is the sequential decision problem?
- What is the state space?
- What are the actions?
- What reward is being optimized?
- What data or interaction is available?
- What algorithm is used?
- What assumptions does the algorithm make?
- How is performance evaluated?
- Why did the method succeed or fail?
Repository
The GitHub repository for the book is:
https://github.com/wanghemath/Book-MathRL
The labs are stored in:
labs/
The final integrated lab is:
labs/chapter-28-lab.ipynb
It can be opened in Google Colab:
https://colab.research.google.com/github/wanghemath/Book-MathRL/blob/main/labs/chapter-28-lab.ipynb
What counts as a project?
A project may be one of the following:
Algorithm comparison project
Compare several reinforcement-learning algorithms on one environment.Environment-design project
Modify or create an environment and study how different algorithms behave.Theory-to-code project
Start from an equation, implement the algorithm, and analyze convergence or failure.Modern RL topic project
Study offline RL, imitation learning, multi-agent RL, model-based RL, or RLHF.Applied modeling project
Formulate an applied problem as an MDP or contextual decision problem and solve a simplified version.Final integrated project
Use Lab 28 as the base and extend it in a meaningful direction.
Core project structure
Every project should include the following parts.
1. Problem formulation
Describe the sequential decision problem.
Include:
- state space \(S\),
- action space \(A\),
- reward function \(R\),
- transition structure,
- discount factor \(\gamma\),
- terminal states, if any,
- policy class.
A finite MDP can be summarized as:
\[ (S,A,P,R,\gamma). \]
A policy is a rule:
\[ \pi(a\mid s). \]
The value of a policy is:
\[ V^\pi(s) = E_\pi \left[ \sum_{t=0}^{\infty} \gamma^t R_{t+1} \mid S_0=s \right]. \]
2. Algorithm description
State the method clearly.
For example, Q-learning uses:
\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma\max_b Q(s',b)-Q(s,a) \right]. \]
SARSA uses:
\[ Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r+\gamma Q(s',a')-Q(s,a) \right]. \]
REINFORCE uses:
\[ \theta \leftarrow \theta + \alpha G_t \nabla_\theta \log \pi_\theta(A_t\mid S_t). \]
Dyna-Q combines real updates with simulated model-based updates.
Offline fitted Q iteration uses a fixed dataset:
\[ D=\{(s_i,a_i,r_i,s_i',d_i)\}_{i=1}^N. \]
3. Experimental design
Report all important experimental settings:
- number of episodes,
- number of random seeds,
- learning rate,
- discount factor,
- exploration schedule,
- function approximation features,
- neural-network size, if used,
- number of planning steps, if model-based,
- offline dataset size, if offline,
- number of expert demonstrations, if imitation learning,
- evaluation episodes.
A project should be reproducible. If another student reads your report, they should be able to rerun the experiment.
4. Evaluation metrics
Choose metrics appropriate for the task.
Common metrics include:
| Metric | Meaning |
|---|---|
| Mean return | Average discounted or undiscounted return |
| Success rate | Fraction of episodes reaching the goal |
| Failure/trap rate | Fraction of episodes reaching bad terminal states |
| Episode length | Number of steps before termination |
| Regret | Loss relative to the best action or policy |
| Bellman error | Size of Bellman residual or TD error |
| Policy entropy | Randomness of a stochastic policy |
| KL divergence | Distance from a reference policy |
| Coverage | Number of visited state-action pairs |
| Model error | Difference between true and learned transition models |
| Preference accuracy | Accuracy of a reward model on comparisons |
5. Visualization
Include at least two meaningful visualizations.
Good options include:
- learning curve,
- rolling average return,
- success-rate curve,
- policy arrow plot,
- value-function heatmap,
- state-visitation heatmap,
- regret curve,
- action-probability plot,
- reward-model scatter plot,
- offline coverage plot,
- comparison bar chart.
A plot should not appear alone. Explain what it shows.
6. Diagnosis
A good project explains why the results happened.
Possible diagnostic themes:
- insufficient exploration,
- large variance,
- biased value estimates,
- bootstrapping instability,
- poor step-size choice,
- bad features,
- distribution shift,
- poor offline coverage,
- model error,
- reward misspecification,
- nonstationarity from other agents,
- reward hacking,
- excessive or insufficient regularization.
7. Conclusion
The conclusion should answer:
- Which method worked best?
- Why did it work best?
- What failed?
- What would you try next?
- What did the experiment teach about reinforcement learning?
Project proposal template
Before starting a final project, write a short proposal.
Title
Give the project a clear title.
Example:
Comparing Q-Learning, Dyna-Q, and Offline FQI in a Stochastic Grid-World
Question
State the main research question.
Examples:
- Does planning improve sample efficiency in this environment?
- How does offline dataset coverage affect learned policy quality?
- Does entropy regularization improve exploration?
- How sensitive is Q-learning to the learning rate?
- Can behavior cloning match expert performance with limited demonstrations?
Environment
Describe the environment:
- state space,
- action space,
- reward,
- transition randomness,
- terminal states.
Methods
List the methods you will compare.
Examples:
- value iteration,
- Q-learning,
- SARSA,
- Dyna-Q,
- REINFORCE,
- actor-critic,
- offline fitted Q iteration,
- behavior cloning.
Metrics
List the metrics you will report.
Examples:
- mean return,
- success rate,
- trap rate,
- episode length,
- coverage,
- policy disagreement,
- regret.
Expected result
State your hypothesis.
Example:
I expect Dyna-Q to learn faster than Q-learning because it uses planning updates from a learned model.
Extension
Describe what is new beyond the base lab.
Examples:
- changed environment,
- new hyperparameter study,
- new diagnostic plot,
- new algorithm variant,
- new comparison table,
- robustness experiment.
Final report template
Use the following structure for the final write-up.
1. Abstract
Write a short summary of the project in 5–8 sentences.
Include:
- the problem,
- the algorithms,
- the main result,
- the main conclusion.
2. Introduction
Explain the motivation.
Example:
This project studies how model-based planning changes sample efficiency in a stochastic navigation task.
3. MDP formulation
Define the MDP:
\[ (S,A,P,R,\gamma). \]
Include:
- states,
- actions,
- rewards,
- transitions,
- terminal states,
- discount factor.
4. Algorithms
Describe each algorithm.
For each method, include:
- the main update equation,
- the data used,
- the policy used for action selection,
- any important hyperparameters.
5. Experimental setup
Report:
- number of training episodes,
- number of evaluation episodes,
- random seeds,
- learning rates,
- exploration schedule,
- planning steps,
- offline dataset size,
- demonstration size.
6. Results
Include:
- a comparison table,
- at least one learning curve,
- at least one policy or value visualization,
- at least one diagnostic plot.
7. Discussion
Explain the results.
Address:
- why some methods worked better,
- why some methods failed,
- whether the results match your hypothesis,
- what limitations remain.
8. Conclusion
Summarize the main lesson.
9. Appendix
Optional. Include extra plots, code snippets, or additional experiments.
Suggested project topics
Topic 1: Q-learning versus SARSA
Question
How do on-policy and off-policy TD control methods differ in risky environments?
Base labs
- Lab 5: Temporal-Difference Learning
- Lab 6: SARSA and Q-Learning in Depth
Suggested experiment
Use a cliff-walking or trap-navigation environment.
Compare:
- SARSA,
- Q-learning,
- Expected SARSA.
Report:
- learning curves,
- final policies,
- average return,
- failure rate.
Key concept
SARSA accounts for exploratory behavior, while Q-learning learns a greedy target policy.
Topic 2: Exploration strategies
Question
Which exploration strategy works best in a sparse-reward environment?
Base labs
- Lab 7: Exploration and Exploitation
- Lab 8: Multi-Armed Bandits
Suggested experiment
Compare:
- fixed epsilon-greedy,
- decaying epsilon-greedy,
- optimistic initialization,
- softmax exploration,
- UCB.
Report:
- cumulative reward,
- regret,
- optimal-action frequency,
- sensitivity to parameters.
Key concept
Exploration is a statistical problem and a control problem at the same time.
Topic 3: Step-size schedules and stochastic approximation
Question
How does the learning-rate schedule affect convergence?
Base labs
- Lab 9: Stochastic Approximation
Suggested experiment
Compare:
\[ \alpha_t = 0.1, \qquad \alpha_t = \frac{1}{t}, \qquad \alpha_t = \frac{1}{\sqrt{t}}, \qquad \alpha_t = \frac{0.5}{10+t}. \]
Report:
- convergence curves,
- variance across seeds,
- final error.
Key concept
Step sizes control the tradeoff between adaptation and convergence.
Topic 4: Function approximation
Question
How do features affect value-function approximation?
Base labs
- Lab 10: Function Approximation
- Lab 11: Linear Value Function Approximation
Suggested experiment
Compare:
- state aggregation,
- polynomial features,
- radial basis functions,
- Fourier features.
Report:
- approximation error,
- learning curves,
- final policy quality,
- feature sensitivity.
Key concept
Function approximation introduces generalization and approximation error.
Topic 5: Policy-gradient variance reduction
Question
How much do reward-to-go and baselines reduce policy-gradient variance?
Base labs
- Lab 12: Policy Gradient Methods
- Lab 13: REINFORCE
- Lab 14: Actor-Critic Methods
Suggested experiment
Compare:
- full-return REINFORCE,
- reward-to-go REINFORCE,
- REINFORCE with baseline,
- actor-critic.
Report:
- learning curves,
- variance across seeds,
- final success rate,
- policy entropy.
Key concept
Baselines reduce variance without changing the expected policy gradient.
Topic 6: Entropy regularization
Question
Does entropy regularization improve exploration and prevent premature policy collapse?
Base labs
- Lab 16: Entropy-Regularized RL
Suggested experiment
Train policies with different entropy temperatures:
\[ \tau = 0,\quad 0.01,\quad 0.05,\quad 0.1,\quad 0.5. \]
Report:
- return,
- entropy,
- action probabilities,
- final policy.
Key concept
Entropy regularization encourages stochasticity and can improve exploration.
Topic 7: DQN-style stabilizers
Question
How do replay buffers and target networks stabilize neural Q-learning?
Base labs
- Lab 18: Deep Q-Networks
Suggested experiment
Compare:
- no replay buffer,
- replay buffer,
- target network,
- replay buffer plus target network.
Report:
- TD loss,
- return,
- final policy,
- instability or divergence.
Key concept
Deep RL requires stabilization because bootstrapping and function approximation can interact badly.
Topic 8: PPO-style clipping
Question
How does clipping affect policy-gradient stability?
Base labs
- Lab 20: Trust Region and PPO-Style Methods
Suggested experiment
Compare clipping parameters:
\[ \epsilon_{\text{clip}} = 0.05,\quad 0.1,\quad 0.2,\quad 0.4. \]
Report:
- return,
- KL divergence,
- clipping fraction,
- policy entropy.
Key concept
PPO-style clipping limits destructive policy updates.
Topic 9: Continuous control
Question
How does Gaussian exploration noise affect continuous control?
Base labs
- Lab 21: Continuous Control
Suggested experiment
Compare Gaussian standard deviations:
\[ \sigma = 0.1,\quad 0.3,\quad 0.6,\quad 1.0. \]
Report:
- return,
- state magnitude,
- action magnitude,
- learned mean controller.
Key concept
Continuous control requires balancing exploration noise with stable control.
Topic 10: Model-based RL
Question
How much data is needed to learn a useful model?
Base labs
- Lab 22: Model-Based Reinforcement Learning
Suggested experiment
Collect datasets of different sizes. Estimate transition models and plan inside them.
Report:
- model error,
- policy success rate,
- state-action coverage,
- learned policy plots.
Key concept
Planning is only as good as the learned model.
Topic 11: Dyna-Q and planning steps
Question
How many planning steps are useful?
Base labs
- Lab 23: Planning and Learning
Suggested experiment
Compare:
\[ n_{\text{planning}} = 0,\quad 1,\quad 5,\quad 20,\quad 50. \]
Report:
- learning curves,
- sample efficiency,
- computation cost,
- final policy.
Key concept
Planning reuses experience but may waste computation or amplify model errors.
Topic 12: Offline RL and coverage
Question
How does offline dataset coverage affect learned policy quality?
Base labs
- Lab 24: Offline Reinforcement Learning
Suggested experiment
Generate:
- random dataset,
- medium-quality dataset,
- expert dataset.
Compare:
- naive offline FQI,
- supported-action FQI,
- count-penalized FQI.
Report:
- coverage,
- learned policies,
- return,
- out-of-distribution action diagnostics.
Key concept
Offline RL is fundamentally limited by dataset support.
Topic 13: Imitation learning and DAgger
Question
How does DAgger reduce distribution shift in behavior cloning?
Base labs
- Lab 25: Imitation Learning
Suggested experiment
Compare:
- behavior cloning with few demonstrations,
- behavior cloning with many demonstrations,
- DAgger-style dataset aggregation.
Report:
- state coverage,
- success rate,
- final policy,
- learner-visited unseen states.
Key concept
DAgger labels states induced by the learner, not only states visited by the expert.
Topic 14: Multi-agent coordination
Question
How stable is independent Q-learning in multi-agent games?
Base labs
- Lab 26: Multi-Agent Reinforcement Learning
Suggested experiment
Compare learning in:
- coordination game,
- matching pennies,
- prisoner’s dilemma,
- cooperative grid task.
Report:
- action frequencies,
- rewards,
- equilibrium behavior,
- seed sensitivity.
Key concept
Other learning agents make the environment nonstationary.
Topic 15: RLHF reward modeling
Question
How does preference-data quality affect reward-model optimization?
Base labs
- Lab 27: Reinforcement Learning from Human Feedback
Suggested experiment
Compare:
- small versus large preference datasets,
- unbiased versus biased preference data,
- weak versus strong KL regularization,
- RLHF-style reward optimization versus DPO-style updates.
Report:
- reward-model rank correlation,
- KL versus reward plots,
- action distributions,
- reward hacking diagnostics.
Key concept
Optimizing a learned reward can fail when the reward model is biased or inaccurate.
Topic 16: Final integrated comparison
Question
Which family of methods works best in a stochastic navigation task?
Base lab
- Lab 28: Final Integrated Reinforcement Learning Project
Suggested experiment
Compare:
- exact dynamic programming,
- Q-learning,
- Dyna-Q,
- offline fitted Q iteration,
- conservative offline learning,
- behavior cloning.
Report:
- final policy plots,
- learning curves,
- success rates,
- coverage diagnostics,
- policy disagreement with the dynamic-programming benchmark.
Key concept
Different RL methods use different information: known models, interaction, offline data, demonstrations, or learned models.
Reproducibility checklist
Before submitting a project, check the following.
Suggested grading rubric
| Category | Excellent | Satisfactory | Needs improvement |
|---|---|---|---|
| MDP formulation | State, action, reward, transition, and objective are clear | Most elements are clear | Formulation is incomplete |
| Mathematical explanation | Correct equations and interpretation | Some equations included | Little connection to math |
| Implementation | Correct, readable, reproducible | Mostly correct | Does not run or is unclear |
| Experimental design | Fair comparison with clear settings | Basic comparison | Missing controls or settings |
| Evaluation | Multiple metrics and meaningful plots | Some metrics | Weak or missing evaluation |
| Diagnosis | Explains why results occurred | Some interpretation | Mostly descriptive |
| Extension | Meaningful original modification | Small modification | No meaningful extension |
| Communication | Clear report with strong organization | Understandable report | Difficult to follow |
Common project mistakes
Avoid the following.
Mistake 1: Only reporting one run
RL algorithms are stochastic. Use multiple seeds when possible.
Mistake 2: Comparing methods unfairly
Use the same environment and evaluation function for all methods.
Mistake 3: Confusing training and evaluation
Exploration is often used during training. Evaluation should usually use the learned greedy or mean policy.
Mistake 4: Ignoring coverage
Offline RL and imitation learning depend heavily on what states and actions appear in the data.
Mistake 5: Showing plots without explanation
Every plot should have a short interpretation.
Mistake 6: Reporting success without diagnosis
A good project explains why the method worked.
Mistake 7: Changing too many things at once
Change one variable at a time when possible.
Mini-project format
For smaller assignments, use this shorter format.
Mini-project report
- Question: What are you testing?
- Method: What algorithm did you use?
- Equation: What is the key update or objective?
- Experiment: What parameter or environment feature did you change?
- Result: What plot or table summarizes the result?
- Interpretation: Why did this happen?
- Conclusion: What did you learn?
A mini-project can be 1–2 pages plus code.
Final project format
For the final project, use a longer format.
Recommended length:
- 4–8 pages,
- plus notebook,
- plus figures and tables.
Required components:
- abstract,
- MDP formulation,
- algorithm descriptions,
- experimental setup,
- results,
- diagnosis,
- conclusion,
- reproducibility notes.
Final project presentation
A short presentation may use the following structure.
Slide 1: Title and question
State the project question.
Slide 2: Environment
Show the state/action/reward setup.
Slide 3: Algorithms
List the methods compared.
Slide 4: Main equations
Show the key update equations.
Slide 5: Results table
Compare final metrics.
Slide 6: Learning curves
Show training behavior.
Slide 7: Diagnostics
Show coverage, model error, policy disagreement, or reward-model issue.
Slide 8: Conclusion
State the main lesson.
A strong final conclusion
A strong conclusion should be specific.
Weak conclusion:
Dyna-Q worked well.
Stronger conclusion:
Dyna-Q reached a high success rate in fewer episodes than Q-learning because planning updates reused each real transition several times. However, its advantage depended on the learned model being accurate; in the changed-environment experiment, outdated model transitions slowed adaptation.
Closing advice
A reinforcement-learning project is successful when it teaches a clear lesson.
The best projects are not necessarily the most complicated. A simple environment with careful comparisons and thoughtful diagnosis is often better than a large experiment with unclear conclusions.
Focus on the mathematical story:
- What is the objective?
- What information does the algorithm use?
- What approximation is made?
- What evidence shows success?
- What evidence explains failure?
That is the core of doing reinforcement learning well.