References

Bibliography for The Mathematics of Reinforcement Learning

References

This page lists references for The Mathematics of Reinforcement Learning. The bibliography includes foundational sources on Markov decision processes, dynamic programming, temporal-difference learning, policy gradients, stochastic approximation, bandits, function approximation, deep reinforcement learning, model-based reinforcement learning, offline reinforcement learning, imitation learning, multi-agent reinforcement learning, and reinforcement learning from human feedback.

Foundational reinforcement learning texts

Core references include Sutton and Barto (2018), Puterman (1994), Bertsekas (2012), Bertsekas and Tsitsiklis (1996a), Bertsekas (2019), Szepesvári (2010), and Powell (2022).

Dynamic programming and Markov decision processes

The mathematical foundation of the book is based on the Bellman equation, dynamic programming, Markov decision processes, contraction mappings, and stochastic stability. Key references include Bellman (1957a), Bellman (1957b), Howard (1960), Blackwell (1962), Banach (1922), Meyn and Tweedie (2009), and Norris (1998).

Stochastic approximation and temporal-difference learning

The stochastic approximation viewpoint is based on Robbins and Monro (1951), Kushner and Yin (2003), and Borkar (2008). Classical TD and control methods are connected to Watkins (1989), Watkins and Dayan (1992), Rummery and Niranjan (1994), Singh and Sutton (1996), Jaakkola et al. (1994), and Tsitsiklis (1994).

Function approximation and approximate dynamic programming

Linear value-function approximation, projection, least-squares TD, and approximate policy iteration are represented by Bradtke and Barto (1996), Boyan (2002), Lagoudakis and Parr (2003), Bertsekas and Tsitsiklis (1996a), Munos (2003), and Munos and Szepesvári (2008).

Policy gradients and actor-critic methods

Policy-gradient and actor-critic chapters draw from Williams (1992), Sutton et al. (2000), Konda and Tsitsiklis (2000), Kakade (2001), Peters and Schaal (2008), Bhatnagar et al. (2009), Schulman et al. (2015), Schulman et al. (2016), and Schulman et al. (2017).

Deep reinforcement learning

The deep RL chapters are informed by Mnih et al. (2015), Mnih et al. (2016), Hasselt et al. (2016), Wang et al. (2016), Schaul et al. (2016), Hessel et al. (2018), Lillicrap et al. (2016), and Haarnoja et al. (2018).

Bandits and exploration

Bandit algorithms and exploration ideas are connected to Thompson (1933), Lai and Robbins (1985), Auer et al. (2002), Russo et al. (2018), and Lattimore and Szepesvári (2020).

Model-based reinforcement learning and planning

Model-based RL, planning, Dyna, prioritized sweeping, and tree-search ideas draw from Sutton (1991), Moore and Atkeson (1993), Deisenroth and Rasmussen (2011), Chua et al. (2018), Coulom (2006), Kocsis and Szepesvári (2006), and Silver et al. (2016).

Offline reinforcement learning

Offline RL and batch RL chapters draw from Lange et al. (2012), Ernst et al. (2005), Fujimoto et al. (2019), Kumar et al. (2020), Agarwal et al. (2020), and Levine et al. (2020).

Imitation learning and inverse reinforcement learning

Imitation learning references include Pomerleau (1989), Ng and Russell (2000), Abbeel and Ng (2004), Ziebart et al. (2008), Ross et al. (2011), and Ho and Ermon (2016).

Multi-agent reinforcement learning and game theory

Multi-agent RL and game-theoretic background draw from Nash (1950), Neumann (1928), Neumann and Morgenstern (1944), Fudenberg and Tirole (1991), Osborne and Rubinstein (1994), Littman (1994), Bowling and Veloso (2002), Hu and Wellman (2003), Lowe et al. (2017), and Foerster et al. (2018).

Reinforcement learning from human feedback

Preference modeling, reward modeling, and RLHF chapters draw from Bradley and Terry (1952), Luce (1959), Christiano et al. (2017), Ziegler et al. (2019), Stiennon et al. (2020), Ouyang et al. (2022), and Rafailov et al. (2023).

General mathematical and machine-learning background

General background references include Goodfellow et al. (2016), Bishop (2006), Hastie et al. (2009), Boyd and Vandenberghe (2004), Nocedal and Wright (2006), Cover and Thomas (2006), and MacKay (2003).

Full bibliography

Abbeel, Pieter, and Andrew Y. Ng. 2004. “Apprenticeship Learning via Inverse Reinforcement Learning.” Proceedings of the Twenty-First International Conference on Machine Learning.
Agarwal, Alekh, Nan Jiang, Sham M. Kakade, and Wen Sun. 2019. “Reinforcement Learning: Theory and Algorithms.” Lecture Notes.
Agarwal, Rishabh, Dale Schuurmans, and Mohammad Norouzi. 2020. “An Optimistic Perspective on Offline Reinforcement Learning.” Proceedings of the 37th International Conference on Machine Learning, 104–14.
Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016. “Concrete Problems in AI Safety.” arXiv Preprint arXiv:1606.06565.
Auer, Peter, Nicolò Cesa-Bianchi, and Paul Fischer. 2002. “Finite-Time Analysis of the Multiarmed Bandit Problem.” Machine Learning 47: 235–56.
Azar, Mohammad Gheshlaghi, Rémi Munos, and Hilbert J. Kappen. 2012. “On the Sample Complexity of Reinforcement Learning with a Generative Model.” Proceedings of the 29th International Conference on Machine Learning.
Baird, Leemon. 1995. “Residual Algorithms: Reinforcement Learning with Function Approximation.” Proceedings of the Twelfth International Conference on Machine Learning, 30–37.
Banach, Stefan. 1922. “Sur Les Opérations Dans Les Ensembles Abstraits Et Leur Application Aux équations Intégrales.” Fundamenta Mathematicae 3: 133–81.
Bellemare, Marc G., Will Dabney, and Rémi Munos. 2017. “A Distributional Perspective on Reinforcement Learning.” Proceedings of the 34th International Conference on Machine Learning, 449–58.
Bellman, Richard. 1957a. “A Markovian Decision Process.” Journal of Mathematics and Mechanics 6 (5): 679–84.
Bellman, Richard. 1957b. Dynamic Programming. Princeton University Press.
Bertsekas, Dimitri P. 2012. Dynamic Programming and Optimal Control. 4th ed. Athena Scientific.
Bertsekas, Dimitri P. 2019. Reinforcement Learning and Optimal Control. Athena Scientific.
Bertsekas, Dimitri P., and John N. Tsitsiklis. 1996a. Neuro-Dynamic Programming. Athena Scientific.
Bertsekas, Dimitri P., and John N. Tsitsiklis. 1996b. “Neuro-Dynamic Programming: An Overview.” Proceedings of the IEEE Conference on Decision and Control.
Bhatnagar, Shalabh, Richard S. Sutton, Mohammad Ghavamzadeh, and Mark Lee. 2009. “Natural Actor-Critic Algorithms.” Automatica 45 (11): 2471–82.
Bishop, Christopher M. 2006. Pattern Recognition and Machine Learning. Springer.
Blackwell, David. 1962. “Discrete Dynamic Programming.” Annals of Mathematical Statistics 33 (2): 719–26.
Borkar, Vivek S. 2008. Stochastic Approximation: A Dynamical Systems Viewpoint. Cambridge University Press.
Bowling, Michael, and Manuela Veloso. 2002. “Multiagent Learning Using a Variable Learning Rate.” Artificial Intelligence 136 (2): 215–50.
Boyan, Justin A. 2002. “Technical Update: Least-Squares Temporal Difference Learning.” Machine Learning 49: 233–46.
Boyd, Stephen, and Lieven Vandenberghe. 2004. Convex Optimization. Cambridge University Press.
Bradley, Ralph Allan, and Milton E. Terry. 1952. “Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons.” Biometrika 39 (3–4): 324–45.
Bradtke, Steven J., and Andrew G. Barto. 1996. “Linear Least-Squares Algorithms for Temporal Difference Learning.” Machine Learning 22: 33–57.
Brémaud, Pierre. 2020. Markov Chains: Gibbs Fields, Monte Carlo Simulation, and Queues. 2nd ed. Springer.
Christiano, Paul F., Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. “Deep Reinforcement Learning from Human Preferences.” Advances in Neural Information Processing Systems 30.
Chua, Kurtland, Roberto Calandra, Rowan McAllister, and Sergey Levine. 2018. “Deep Reinforcement Learning in a Handful of Trials Using Probabilistic Dynamics Models.” Advances in Neural Information Processing Systems 31.
Coulom, Rémi. 2006. “Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search.” Computers and Games, 72–83.
Cover, Thomas M., and Joy A. Thomas. 2006. Elements of Information Theory. 2nd ed. Wiley.
Deisenroth, Marc Peter, and Carl Edward Rasmussen. 2011. “PILCO: A Model-Based and Data-Efficient Approach to Policy Search.” Proceedings of the 28th International Conference on Machine Learning, 465–72.
Ernst, Damien, Pierre Geurts, and Louis Wehenkel. 2005. “Tree-Based Batch Mode Reinforcement Learning.” Journal of Machine Learning Research 6: 503–56.
Farahmand, Amir-massoud, Mohammad Ghavamzadeh, Csaba Szepesvári, and Rémi Munos. 2010. “Regularized Policy Iteration.” Advances in Neural Information Processing Systems 23.
Foerster, Jakob N., Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018. “Counterfactual Multi-Agent Policy Gradients.” Proceedings of the AAAI Conference on Artificial Intelligence 32 (1).
Fudenberg, Drew, and Jean Tirole. 1991. Game Theory. MIT Press.
Fujimoto, Scott, David Meger, and Doina Precup. 2019. “Off-Policy Deep Reinforcement Learning Without Exploration.” Proceedings of the 36th International Conference on Machine Learning, 2052–62.
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press.
Haarnoja, Tuomas, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.” Proceedings of the 35th International Conference on Machine Learning, 1861–70.
Hasselt, Hado van, Arthur Guez, and David Silver. 2016. “Deep Reinforcement Learning with Double q-Learning.” Proceedings of the AAAI Conference on Artificial Intelligence 30 (1).
Hastie, Trevor, Robert Tibshirani, and Jerome Friedman. 2009. The Elements of Statistical Learning. 2nd ed. Springer.
Hessel, Matteo, Joseph Modayil, Hado van Hasselt, et al. 2018. “Rainbow: Combining Improvements in Deep Reinforcement Learning.” Proceedings of the AAAI Conference on Artificial Intelligence 32 (1).
Ho, Jonathan, and Stefano Ermon. 2016. “Generative Adversarial Imitation Learning.” Advances in Neural Information Processing Systems 29.
Howard, Ronald A. 1960. “Dynamic Programming and Markov Processes.” Technology Press and Wiley.
Hu, Junling, and Michael P. Wellman. 2003. “Nash q-Learning for General-Sum Stochastic Games.” Journal of Machine Learning Research 4: 1039–69.
Jaakkola, Tommi, Michael I. Jordan, and Satinder P. Singh. 1994. “On the Convergence of Stochastic Iterative Dynamic Programming Algorithms.” Neural Computation 6 (6): 1185–201.
Jaques, Natasha, Asma Ghandeharioun, Judy Hanwen Shen, et al. 2019. “Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.” arXiv Preprint arXiv:1907.00456.
Jin, Chi, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan. 2018. “Is q-Learning Provably Efficient?” Advances in Neural Information Processing Systems 31.
Kakade, Sham M. 2001. “A Natural Policy Gradient.” Advances in Neural Information Processing Systems 14.
Kochenderfer, Mykel J., Tim A. Wheeler, and Kyle H. Wray. 2022. Algorithms for Decision Making. MIT Press.
Kocsis, Levente, and Csaba Szepesvári. 2006. “Bandit Based Monte-Carlo Planning.” European Conference on Machine Learning, 282–93.
Konda, Vijay R., and John N. Tsitsiklis. 2000. “Actor-Critic Algorithms.” Advances in Neural Information Processing Systems 12.
Kumar, Aviral, Aurick Zhou, George Tucker, and Sergey Levine. 2020. “Conservative q-Learning for Offline Reinforcement Learning.” Advances in Neural Information Processing Systems 33: 1179–91.
Kushner, Harold J., and Dean S. Clark. 1978. “Stochastic Approximation Methods for Constrained and Unconstrained Systems.” Applied Mathematical Sciences 26.
Kushner, Harold J., and G. George Yin. 2003. Stochastic Approximation and Recursive Algorithms and Applications. 2nd ed. Springer.
Lagoudakis, Michail G., and Ronald Parr. 2003. “Least-Squares Policy Iteration.” Journal of Machine Learning Research 4: 1107–49.
Lai, Tze Leung, and Herbert Robbins. 1985. “Asymptotically Efficient Adaptive Allocation Rules.” Advances in Applied Mathematics 6 (1): 4–22.
Lange, Sascha, Thomas Gabel, and Martin Riedmiller. 2012. “Batch Reinforcement Learning.” Reinforcement Learning: State-of-the-Art, 45–73.
Lattimore, Tor, and Csaba Szepesvári. 2020. Bandit Algorithms. Cambridge University Press.
Levine, Sergey, Aviral Kumar, George Tucker, and Justin Fu. 2020. “Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.” arXiv Preprint arXiv:2005.01643.
Lillicrap, Timothy P., Jonathan J. Hunt, Alexander Pritzel, et al. 2016. “Continuous Control with Deep Reinforcement Learning.” International Conference on Learning Representations.
Littman, Michael L. 1994. “Markov Games as a Framework for Multi-Agent Reinforcement Learning.” Machine Learning Proceedings, 157–63.
Lowe, Ryan, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017. “Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments.” Advances in Neural Information Processing Systems 30.
Luce, R. Duncan. 1959. “Individual Choice Behavior: A Theoretical Analysis.” Wiley.
MacKay, David J. C. 2003. Information Theory, Inference, and Learning Algorithms. Cambridge University Press.
Meyn, Sean P., and Richard L. Tweedie. 2009. Markov Chains and Stochastic Stability. 2nd ed. Cambridge University Press.
Mnih, Volodymyr, Adrià Puigdomènech Badia, Mehdi Mirza, et al. 2016. “Asynchronous Methods for Deep Reinforcement Learning.” Proceedings of the 33rd International Conference on Machine Learning, 1928–37.
Mnih, Volodymyr, Koray Kavukcuoglu, David Silver, et al. 2015. “Human-Level Control Through Deep Reinforcement Learning.” Nature 518: 529–33.
Moore, Andrew W., and Christopher G. Atkeson. 1993. “Prioritized Sweeping: Reinforcement Learning with Less Data and Less Time.” Machine Learning 13: 103–30.
Munos, Rémi. 2003. “Error Bounds for Approximate Policy Iteration.” Proceedings of the 20th International Conference on Machine Learning, 560–67.
Munos, Rémi, and Csaba Szepesvári. 2008. “Finite-Time Bounds for Fitted Value Iteration.” Journal of Machine Learning Research 9: 815–57.
Nachum, Ofir, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans. 2017. “Bridging the Gap Between Value and Policy Based Reinforcement Learning.” Advances in Neural Information Processing Systems 30.
Nash, John F. 1950. “Equilibrium Points in n-Person Games.” Proceedings of the National Academy of Sciences 36 (1): 48–49.
Neumann, John von. 1928. “Zur Theorie Der Gesellschaftsspiele.” Mathematische Annalen 100: 295–320.
Neumann, John von, and Oskar Morgenstern. 1944. Theory of Games and Economic Behavior. Princeton University Press.
Ng, Andrew Y., Daishi Harada, and Stuart J. Russell. 1999. “Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping.” Proceedings of the Sixteenth International Conference on Machine Learning, 278–87.
Ng, Andrew Y., and Stuart J. Russell. 2000. “Algorithms for Inverse Reinforcement Learning.” Proceedings of the Seventeenth International Conference on Machine Learning, 663–70.
Nocedal, Jorge, and Stephen J. Wright. 2006. Numerical Optimization. 2nd ed. Springer.
Norris, J. R. 1998. Markov Chains. Cambridge University Press.
Osborne, Martin J., and Ariel Rubinstein. 1994. A Course in Game Theory. MIT Press.
Ouyang, Long, Jeffrey Wu, Xu Jiang, et al. 2022. “Training Language Models to Follow Instructions with Human Feedback.” Advances in Neural Information Processing Systems 35: 27730–44.
Peters, Jan, and Stefan Schaal. 2008. “Natural Actor-Critic.” Neurocomputing 71 (7–9): 1180–90.
Pomerleau, Dean A. 1989. “ALVINN: An Autonomous Land Vehicle in a Neural Network.” Advances in Neural Information Processing Systems 1.
Powell, Warren B. 2022. Reinforcement Learning and Stochastic Optimization: A Unified Framework for Sequential Decisions. Wiley.
Puterman, Martin L. 1994. Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley.
Rafailov, Rafael, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. “Direct Preference Optimization: Your Language Model Is Secretly a Reward Model.” Advances in Neural Information Processing Systems 36.
Robbins, Herbert, and Sutton Monro. 1951. “A Stochastic Approximation Method.” Annals of Mathematical Statistics 22 (3): 400–407.
Ross, Sheldon M. 2019. Introduction to Probability Models. 12th ed. Academic Press.
Ross, Stéphane, Geoffrey J. Gordon, and J. Andrew Bagnell. 2011. “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning.” Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 627–35.
Rummery, Gavin A., and Mahesan Niranjan. 1994. On-Line q-Learning Using Connectionist Systems. CUED/F-INFENG/TR 166. Cambridge University Engineering Department.
Russo, Daniel, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen. 2018. “A Tutorial on Thompson Sampling.” Foundations and Trends in Machine Learning 11 (1): 1–96.
Schaul, Tom, John Quan, Ioannis Antonoglou, and David Silver. 2016. “Prioritized Experience Replay.” International Conference on Learning Representations.
Schulman, John, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel. 2015. “Trust Region Policy Optimization.” Proceedings of the 32nd International Conference on Machine Learning, 1889–97.
Schulman, John, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel. 2016. “High-Dimensional Continuous Control Using Generalized Advantage Estimation.” International Conference on Learning Representations.
Schulman, John, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. “Proximal Policy Optimization Algorithms.” arXiv Preprint arXiv:1707.06347.
Silver, David, Aja Huang, Chris J. Maddison, et al. 2016. “Mastering the Game of Go with Deep Neural Networks and Tree Search.” Nature 529: 484–89.
Singh, Satinder P., and Richard S. Sutton. 1996. “Reinforcement Learning with Replacing Eligibility Traces.” Machine Learning 22: 123–58.
Stiennon, Nisan, Long Ouyang, Jeffrey Wu, et al. 2020. “Learning to Summarize with Human Feedback.” Advances in Neural Information Processing Systems 33: 3008–21.
Sutton, Richard S. 1991. “Dyna, an Integrated Architecture for Learning, Planning, and Reacting.” ACM SIGART Bulletin 2 (4): 160–63.
Sutton, Richard S., and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. 2nd ed. MIT Press.
Sutton, Richard S., David McAllester, Satinder Singh, and Yishay Mansour. 2000. “Policy Gradient Methods for Reinforcement Learning with Function Approximation.” Advances in Neural Information Processing Systems 12.
Szepesvári, Csaba. 2010. Algorithms for Reinforcement Learning. Synthesis Lectures on Artificial Intelligence and Machine Learning. Morgan & Claypool Publishers.
Tesauro, Gerald. 1995. “Temporal Difference Learning and TD-Gammon.” Communications of the ACM 38 (3): 58–68.
Thompson, William R. 1933. “On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples.” Biometrika 25 (3–4): 285–94.
Tsitsiklis, John N. 1994. “Asynchronous Stochastic Approximation and q-Learning.” Machine Learning 16: 185–202.
Wang, Ziyu, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas. 2016. “Dueling Network Architectures for Deep Reinforcement Learning.” Proceedings of the 33rd International Conference on Machine Learning, 1995–2003.
Watkins, Christopher J. C. H. 1989. “Learning from Delayed Rewards.” PhD Thesis, King’s College, University of Cambridge.
Watkins, Christopher J. C. H., and Peter Dayan. 1992. “Q-Learning.” Machine Learning 8 (3–4): 279–92.
Wiering, Marco, and Martijn van Otterlo. 2012. “Reinforcement Learning: State-of-the-Art.” Adaptation, Learning, and Optimization 12.
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning.” Machine Learning 8: 229–56.
Ziebart, Brian D., Andrew L. Maas, J. Andrew Bagnell, and Anind K. Dey. 2008. “Maximum Entropy Inverse Reinforcement Learning.” Proceedings of the AAAI Conference on Artificial Intelligence, 1433–38.
Ziegler, Daniel M., Nisan Stiennon, Jeffrey Wu, et al. 2019. “Fine-Tuning Language Models from Human Preferences.” arXiv Preprint arXiv:1909.08593.