An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance

An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance

ACL ARR 2024 June Submission3418 Authors

16 Jun 2024 (modified: 02 Jul 2024)ACL ARR 2024 June SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Abstract: Given the rise of multimedia content, human translators increasingly focus on culturally adapting not only words but also other modalities such as images to convey the same meaning. While several applications stand to benefit from this, machine translation systems remain confined to dealing with language in speech and text. In this work, we introduce a new task of translating images to make them culturally relevant. First, we build three pipelines comprising state-of-the-art generative models to do the task. Next, we build a two-part evaluation dataset -- (i) concept: comprising 600 images that are cross-culturally coherent, focusing on a single concept per image; and (ii) application: comprising 100 images curated from real-world applications. We conduct a multi-faceted human evaluation of translated images to assess for cultural relevance and meaning preservation. We find that as of today, image-editing models fail at this task, but can be improved by leveraging LLMs and retrievers in the loop. Best pipelines can only translate 5% of images for some countries in the easier concept dataset and no translation is successful for some countries in the application dataset, highlighting the challenging nature of the task. Our code and data is released here: https://anonymous.4open.science/r/image-translation-6980.

Paper Type: Long

Research Area: Resources and Evaluation

Research Area Keywords: corpus creation, human evaluation, cross-modal machine translation

Contribution Types: Approaches to low-resource settings, Data resources

Languages Studied: English

Submission Number: 3418

Loading