Parametrized Quantum Policies for Reinforcement Learning

Sofiene Jerbi; Casper Gyurik; Simon Callum Marshall; Hans J Briegel; Vedran Dunjko

Parametrized Quantum Policies for Reinforcement Learning

Sofiene Jerbi, Casper Gyurik, Simon Callum Marshall, Hans J Briegel, Vedran Dunjko

Published: 09 Nov 2021, Last Modified: 05 May 2023NeurIPS 2021 PosterReaders: Everyone

Keywords: reinforcement learning, quantum computing, parametrized quantum circuits, quantum neural networks, policy gradient, quantum machine learning, quantum reinforcement learning, quantum, variational quantum circuits

TL;DR: We investigate the potential of parametrized quantum computations when trained as reinforcement learning policies in classical environments.

Abstract: With the advent of real-world quantum computing, the idea that parametrized quantum computations can be used as hypothesis families in a quantum-classical machine learning system is gaining increasing traction. Such hybrid systems have already shown the potential to tackle real-world tasks in supervised and generative learning, and recent works have established their provable advantages in special artificial tasks. Yet, in the case of reinforcement learning, which is arguably most challenging and where learning boosts would be extremely valuable, no proposal has been successful in solving even standard benchmarking tasks, nor in showing a theoretical learning advantage over classical algorithms. In this work, we achieve both. We propose a hybrid quantum-classical reinforcement learning model using very few qubits, which we show can be effectively trained to solve several standard benchmarking environments. Moreover, we demonstrate, and formally prove, the ability of parametrized quantum circuits to solve certain learning tasks that are intractable to classical models, including current state-of-art deep neural networks, under the widely-believed classical hardness of the discrete logarithm problem.

Code Of Conduct: I certify that all co-authors of this work have read and commit to adhering to the NeurIPS Statement on Ethics, Fairness, Inclusivity, and Code of Conduct.

Supplementary Material: pdf

Code: https://www.tensorflow.org/quantum/tutorials/quantum_reinforcement_learning

12 Replies

Loading