Benchmarking and optimizing organism wide single-cell RNA alignment methods

Juan Javier Díaz-Mejía; Elias Williams; Octavian Focsa; Dylan Mendonca; Swechha Singh; Brendan Innes; Samuel Cooper

Benchmarking and optimizing organism wide single-cell RNA alignment methods

Juan Javier Díaz-Mejía, Elias Williams, Octavian Focsa, Dylan Mendonca, Swechha Singh, Brendan Innes, Samuel Cooper

Published: 06 Mar 2025, Last Modified: 21 Jul 2025ICLR 2025 Workshop LMRLEveryoneRevisionsBibTeXCC BY 4.0

Track: Full Paper Track

Keywords: single-cell RNA, Benchmark, Adversarial Learning, Variational Inference, Genomics, Embedding, Single-cell

TL;DR: We use two new benchmarks and a new metric to test how well models can align single-cell RNA data at scale

Abstract: Many methods have been proposed for removing batch effects and aligning single-cell RNA (scRNA) datasets. However, performance is typically evaluated based on multiple parameters and few datasets, creating challenges in assessing which method is best for aligning data at scale. Here, we introduce the K-Neighbors Intersection (KNI) score, a single score that both penalizes batch effects and measures accuracy at cross-dataset cell-type label prediction alongside carefully curated small (scMARK) and large (scREF) benchmarks comprising 11 and 46 human scRNA studies respectively, where we have standardized author labels. Using the KNI score, we evaluate and optimize published approaches for cross-dataset single-cell RNA integration. We introduce Batch Adversarial single-cell Variational Inference (BA-scVI), as a new variant of scVI that uses adversarial training to penalize batch-effects in the encoder and decoder, and show this approach outperforms other methods. In the resulting aligned space, we find that the granularity of cell-type groupings is conserved, supporting the notion that whole-organism cell-type maps can be created by a single model without loss of information.

Attendance: Samuel Cooper

Submission Number: 53

Loading