Skew-Explore: Learn faster in continuous spaces with sparse rewards

Xi Chen; Yuan Gao; Ali Ghadirzadeh; Marten Bjorkman; Ginevra Castellano; Patric Jensfelt

Skew-Explore: Learn faster in continuous spaces with sparse rewards

Xi Chen, Yuan Gao, Ali Ghadirzadeh, Marten Bjorkman, Ginevra Castellano, Patric Jensfelt

25 Sept 2019 (modified: 05 May 2023)ICLR 2020 Conference Blind SubmissionReaders: Everyone

Abstract: In many reinforcement learning settings, rewards which are extrinsically available to the learning agent are too sparse to train a suitable policy. Beside reward shaping which requires human expertise, utilizing better exploration strategies helps to circumvent the problem of policy training with sparse rewards. In this work, we introduce an exploration approach based on maximizing the entropy of the visited states while learning a goal-conditioned policy. The main contribution of this work is to introduce a novel reward function which combined with a goal proposing scheme, increases the entropy of the visited states faster compared to the prior work. This improves the exploration capability of the agent, and therefore enhances the agent's chance to solve sparse reward problems more efficiently. Our empirical studies demonstrate the superiority of the proposed method to solve different sparse reward problems in comparison to the prior work.

Code: https://anonymous.4open.science/r/b4596073-4cbc-4ac6-b85b-e9a786909058/

Keywords: reinforcement learning, exploration, sparse reward

Original Pdf: pdf

9 Replies

Loading