Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation

Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation

TMLR Paper1166 Authors

16 May 2023 (modified: 17 Sept 2024)Rejected by TMLREveryoneRevisionsBibTeXCC BY 4.0

Abstract: We introduce Probabilistic Rank and Reward (PRR), a scalable probabilistic model for personalized slate recommendation. Our approach allows off-policy estimation of the reward in the ubiquitous scenario where the user interacts with at most one item from a slate of K items. We show that the probability of a slate being successful can be learned efficiently by combining the reward, whether the user successfully interacted with the slate, and the rank, the item that was selected within the slate. PRR outperforms existing off-policy reward optimizing methods and is far more scalable to large action spaces. Moreover, PRR allows fast delivery of recommendations powered by maximum inner product search (MIPS), making it suitable in low latency domains such as computational advertising.

Submission Length: Regular submission (no more than 12 pages of main content)

Assigned Action Editor: ~Francisco_J._R._Ruiz1

Submission Number: 1166

Loading