2019 (modified: 24 Feb 2022)ICLR (Poster) 2019Readers: Everyone
Abstract:A backward model of previous (state, action) given the next state, i.e. P(s_t, a_t | s_{t+1}), can be used to simulate additional trajectories terminating at states of interest! Improves RL learning efficiency.