Learning Stackelberg Equilibria and Applications to Economic Design Games

Gianluca Brero; Darshan Chakrabarti; Alon Eden; Matthias Gerstgrasser; Vincent Li; David C. Parkes

Learning Stackelberg Equilibria and Applications to Economic Design Games

Gianluca Brero, Darshan Chakrabarti, Alon Eden, Matthias Gerstgrasser, Vincent Li, David C. Parkes

22 Sept 2022 (modified: 13 Feb 2023)ICLR 2023 Conference Withdrawn SubmissionReaders: Everyone

Keywords: Multi-agent Systems, Reinforcement Learning, Economic Design

Abstract: We study the use of reinforcement learning to learn the optimal leader's strategy in Stackelberg games. Learning a leader’s strategy has an innate stationarity problem---when optimizing the leader’s strategy, the followers’ strategies might shift. To circumvent this problem, we model the followers via no-regret dynamics to converge to a Bayesian Coarse-Correlated Equilibrium (B-CCE) of the game induced by the leader. We then embed the followers' no-regret dynamics in the leader's learning environment, which allows us to formulate our learning problem as a standard POMDP. We prove that the optimal policy of this POMDP achieves the same utility as the optimal leader's strategy in our Stackelberg game. We solve this POMDP using actor-critic methods, where the critic is given access to the joint information of all the agents. Finally, we show that our methods are able to learn optimal leader strategies in a variety of settings of increasing complexity, including indirect mechanisms where the leader’s strategy is setting up the mechanism’s rules.

Anonymous Url: I certify that there is no URL (e.g., github page) that could be used to find authors’ identity.

No Acknowledgement Section: I certify that there is no acknowledgement section in this submission for double blind review.

Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics

Submission Guidelines: Yes

Please Choose The Closest Area That Your Submission Falls Into: Reinforcement Learning (eg, decision and control, planning, hierarchical RL, robotics)

Supplementary Material: zip

10 Replies

Loading