Indexed Minimum Empirical Divergence for Unimodal Bandits

Hassan SABER; Pierre MENARD; Odalric-Ambrym Maillard

Indexed Minimum Empirical Divergence for Unimodal Bandits

Hassan SABER, Pierre MENARD, Odalric-Ambrym Maillard

Published: 09 Nov 2021, Last Modified: 05 May 2023NeurIPS 2021 PosterReaders: Everyone

Keywords: Multi-Armed Bandit, Indexed Minimum Empirical Divergence, Unimodal Bandits, Optimal Algorithm, One-Dimensional Exponential Family Distributions, Regret analysis

TL;DR: Optimal strategy for the unimodal bandit problem with one-dimentional family distributions. Elegant proof and improved practical performances.

Abstract: We consider a stochastic multi-armed bandit problem specified by a set of one-dimensional family exponential distributions endowed with a unimodal structure. The unimodal structure is of practical relevance for several applications. We introduce IMED-UB, an algorithm that exploits provably optimally the unimodal-structure, by adapting to this setting the Indexed Minimum Empirical Divergence (IMED) algorithm introduced by Honda and Takemura (2015). Owing to our proof technique, we are able to provide a concise finite-time analysis of the IMED-UB algorithm, that is simple and yet yields asymptotic optimality. We finally provide numerical experiments showing that IMED-UB competes favorably with the recently introduced state-of-the-art algorithms.

Code Of Conduct: I certify that all co-authors of this work have read and commit to adhering to the NeurIPS Statement on Ethics, Fairness, Inclusivity, and Code of Conduct.

Supplementary Material: pdf

8 Replies

Loading