Benchopt: Reproducible, efficient and collaborative optimization benchmarks

Thomas Moreau; Mathurin Massias; Alexandre Gramfort; Pierre Ablin; Pierre-Antoine Bannier; Benjamin Charlier; Mathieu Dagréou; Tom Dupre la Tour; Ghislain Durif; Cássio Fraga Dantas; Quentin Klopfenstein; Johan Larsson; En Lai; Tanguy Lefort; Benoît Malézieux; Badr Moufad; Binh Nguyen; Alain Rakotomamonjy; Zaccharie Ramzi; Joseph Salmon; Samuel Vaiter

Benchopt: Reproducible, efficient and collaborative optimization benchmarks

Published: 31 Oct 2022, Last Modified: 06 Apr 2025NeurIPS 2022 AcceptReaders: Everyone

Keywords: reproducibility, optimization, lasso, resnet, logistic regression, open source software, benchmark

TL;DR: Collaborative framework to automate, publish and reproduce optimization benchmarks in machine learning across programming languages and hardware architectures.

Abstract: Numerical validation is at the core of machine learning research as it allows us to assess the actual impact of new methods, and to confirm the agreement between theory and practice. Yet, the rapid development of the field poses several challenges: researchers are confronted with a profusion of methods to compare, limited transparency and consensus on best practices, as well as tedious re-implementation work. As a result, validation is often very partial, which can lead to wrong conclusions that slow down the progress of research. We propose Benchopt, a collaborative framework to automatize, publish and reproduce optimization benchmarks in machine learning across programming languages and hardware architectures. Benchopt simplifies benchmarking for the community by providing an off-the-shelf tool for running, sharing and extending experiments. To demonstrate its broad usability, we showcase benchmarks on three standard ML tasks: $\ell_2$-regularized logistic regression, Lasso and ResNet18 training for image classification. These benchmarks highlight key practical findings that give a more nuanced view of state-of-the-art for these problems, showing that for practical evaluation, the devil is in the details.

Supplementary Material: zip

Community Implementations: [![CatalyzeX](/images/catalyzex_icon.svg) 10 code implementations](https://www.catalyzex.com/paper/benchopt-reproducible-efficient-and/code)

12 Replies

Loading