Large-scale similarity search with Optimal Transport

Cléa Mehnia Laouar; Yuki Takezawa; Makoto Yamada

Large-scale similarity search with Optimal Transport

Cléa Mehnia Laouar, Yuki Takezawa, Makoto Yamada

Published: 07 Oct 2023, Last Modified: 01 Dec 2023EMNLP 2023 MainEveryoneRevisionsBibTeX

Submission Type: Regular Short Paper

Submission Track: Machine Learning for NLP

Submission Track 2: Efficient Methods for NLP

Keywords: Optimal transport, Wasserstein distance, Document Classification

Abstract: Wasserstein distance is a powerful tool for comparing probability distributions and is widely used for document classification and retrieval tasks in NLP. In particular, it is known as the word mover's distance (WMD) in the NLP community. WMD exhibits excellent performance for various NLP tasks; however, one of its limitations is its computational cost and thus is not useful for large-scale distribution comparisons. In this study, we propose a simple and effective nearest neighbor search based on the Wasserstein distance. Specifically, we employ the L1 embedding method based on the tree-based Wasserstein approximation and subsequently used the nearest neighbor search to efficiently find the $k$-nearest neighbors. Through benchmark experiments, we demonstrate that the proposed approximation has comparable performance to the vanilla Wasserstein distance and can be computed three orders of magnitude faster than the vanilla Wasserstein distance.

Submission Number: 1196

Loading