LRVS-Fashion: Extending Visual Search with Referring Instructions

TMLR Paper4648 Authors

10 Apr 2025 (modified: 13 Jun 2025)Under review for TMLREveryoneRevisionsBibTeXCC BY 4.0
Abstract: This paper introduces a new challenge for image similarity search in the context of fashion, addressing the inherent ambiguity in this domain stemming from complex images. We present Referred Visual Search (RVS), a task allowing users to define more precisely the desired similarity, following recent interest in the industry. We release a new large public dataset, LRVS-Fashion, consisting of 272k fashion products with 842k images extracted from fashion catalogs, designed explicitly for this task. However, unlike traditional visual search methods in the industry, we demonstrate that superior performance can be achieved by bypassing explicit object detection and adopting weakly-supervised conditional contrastive learning on image tuples. Our method is lightweight and demonstrates robustness, reaching Recall at one superior to strong detection-based baselines against 2M distractors.
Submission Length: Regular submission (no more than 12 pages of main content)
Assigned Action Editor: ~Vinay_P_Namboodiri1
Submission Number: 4648
Loading