Multiple Testing for IR and Recommendation System Experiments

Ngozi Ihemelandu, Michael D. Ekstrand

Published: 2024, Last Modified: 28 Apr 2025ECIR (3) 2024EveryoneRevisionsBibTeXCC BY-SA 4.0

Abstract: While there has been significant research on statistical techniques for comparing two information retrieval (IR) systems, many IR experiments test more than two systems. This can lead to inflated false discoveries due to the multiple-comparison problem (MCP). A few IR studies have investigated multiple comparison procedures; these studies mostly use TREC data and control the familywise error rate. In this study, we extend their investigation to include recommendation system evaluation data as well as multiple comparison procedures that controls for False Discovery Rate (FDR).