Compute- and Memory-Efficient Reinforcement Learning with Latent Experience Replay

Lili Chen; Kimin Lee; Aravind Srinivas; Pieter Abbeel

Compute- and Memory-Efficient Reinforcement Learning with Latent Experience Replay

Lili Chen, Kimin Lee, Aravind Srinivas, Pieter Abbeel

28 Sept 2020 (modified: 05 May 2023)ICLR 2021 Conference Blind SubmissionReaders: Everyone

Keywords: reinforcement learning, deep learning, computational efficiency, memory efficiency

Abstract: Recent advances in off-policy deep reinforcement learning (RL) have led to impressive success in complex tasks from visual observations. Experience replay improves sample-efficiency by reusing experiences from the past, and convolutional neural networks (CNNs) process high-dimensional inputs effectively. However, such techniques demand high memory and computational bandwidth. In this paper, we present Latent Vector Experience Replay (LeVER), a simple modification of existing off-policy RL methods, to address these computational and memory requirements without sacrificing the performance of RL agents. To reduce the computational overhead of gradient updates in CNNs, we freeze the lower layers of CNN encoders early in training due to early convergence of their parameters. Additionally, we reduce memory requirements by storing the low-dimensional latent vectors for experience replay instead of high-dimensional images, enabling an adaptive increase in the replay buffer capacity, a useful technique in constrained-memory settings. In our experiments, we show that LeVER does not degrade the performance of RL agents while significantly saving computation and memory across a diverse set of DeepMind Control environments and Atari games. Finally, we show that LeVER is useful for computation-efficient transfer learning in RL because lower layers of CNNs extract generalizable features, which can be used for different tasks and domains.

Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics

One-sentence Summary: We present a compute- and memory-efficient modification of off-policy RL algorithms by freezing lower layers of CNN encoders early in training.

Reviewed Version (pdf): https://openreview.net/references/pdf?id=JxOn09Yv9v

12 Replies

Loading