CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

David Orlando Romero Mogrovejo; Chenyang Lyu; Haryo Akbarianto Wibowo; Santiago Góngora; Aishik Mandal; Sukannya Purkayastha; Jesus-German Ortiz-Barajas; Emilio Villa Cueva; Jinheon Baek; Soyeong Jeong; Injy Hamed; Zheng Xin Yong; Zheng Wei Lim; Paula Mónica Silva; Jocelyn Dunstan; Mélanie Jouitteau; David LE MEUR; Joan Nwatu; Ganzorig Batnasan; Munkh-Erdene Otgonbold; Munkhjargal Gochoo; Guido Ivetta; Luciana Benotti; Laura Alonso Alemany; Hernán Maina; Jiahui Geng; Tiago Timponi Torrent; Frederico Belcavello; Marcelo Viridiano; Jan Christian Blaise Cruz; Dan John Velasco; Oana Ignat; Zara Burzo; Chenxi Whitehouse; Artem Abzaliev; Teresa Clifford; Gráinne Caulfield; Teresa Lynn; Christian Salamea-Palacios; Vladimir Araujo; Yova Kementchedjhieva; Mihail Minkov Mihaylov; Israel Abebe Azime; Henok Biadglign Ademtew; Bontu Fufa Balcha; Naome A Etori; David Ifeoluwa Adelani; Rada Mihalcea; Atnafu Lambebo Tonja; Maria Camila Buitrago Cabrera; Gisela Vallejo; Holy Lovenia; Ruochen Zhang; Marcos Estecha-Garitagoitia; Mario Rodríguez-Cantelar; Toqeer Ehsan; Rendi Chevi; Muhammad Farid Adilazuarda; Ryandito Diandaru; Samuel Cahyawijaya; Fajri Koto; Tatsuki Kuribayashi; Haiyue Song; Aditya Nanda Kishore Khandavally; Thanmay Jayakumar; Raj Dabre; Mohamed Fazli Mohamed Imam; Kumaranage Ravindu Yasas Nagasinghe; Alina Dragonetti; Luis Fernando D'Haro; Olivier NIYOMUGISHA; Jay Gala; Pranjal A Chitale; Fauzan Farooqui; Thamar Solorio; Alham Fikri Aji

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Published: 26 Sept 2024, Last Modified: 13 Nov 2024NeurIPS 2024 Track Datasets and Benchmarks OralEveryoneRevisionsBibTeXCC BY-SA 4.0

Keywords: Multimodality, Multicultural, Multilingual, VQA, Dataset, Benchmark

Abstract: Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field.

Supplementary Material: pdf

Flagged For Ethics Review: true

Submission Number: 2414

Loading