A Large Deviations Perspective on Policy Gradient Algorithms

Wouter Jongeneel, Daniel Kuhn, Mengmeng Li

Published: 01 May 2024, Last Modified: 25 May 2024L4DCEveryoneCC BY 4.0

Abstract: Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent for possibly nonconvex objectives satisfying a Polyak-Łojasiewicz condition. Leveraging the contraction principle from large deviations theory, we illustrate the potential of this result by showing how convergence properties of policy gradient with a softmax parametrization and an entropy regularized objective can be naturally extended to a wide spectrum of other policy parametrizations.