Recurrent Reinforcement Learning with Memoroids

Steven Morad; Chris Lu; Ryan Kortvelesy; Stephan Liwicki; Jakob Nicolaus Foerster; Amanda Prorok

Recurrent Reinforcement Learning with Memoroids

Steven Morad, Chris Lu, Ryan Kortvelesy, Stephan Liwicki, Jakob Nicolaus Foerster, Amanda Prorok

Published: 25 Sept 2024, Last Modified: 06 Nov 2024NeurIPS 2024 posterEveryoneRevisionsBibTeXCC BY 4.0

Keywords: POMDP, reinforcement learning, memory models, recurrent neural network

TL;DR: We propose a new memory model framework and batching method to improve time, space, and sample efficiency in POMDPs

Abstract: Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov states. Neither model scales particularly well to long sequences, especially compared to an emerging class of memory models called Linear Recurrent Models. We discover that the recurrent update of these models resembles a monoid, leading us to reformulate existing models using a novel monoid-based framework that we call memoroids. We revisit the traditional approach to batching in recurrent reinforcement learning, highlighting theoretical and empirical deficiencies. We leverage memoroids to propose a batching method that improves sample efficiency, increases the return, and simplifies the implementation of recurrent loss functions in reinforcement learning.

Supplementary Material: zip

Primary Area: Reinforcement learning

Submission Number: 5217

Loading