MemeCap: A Dataset for Captioning and Interpreting Memes

EunJeong Hwang; Vered Shwartz

MemeCap: A Dataset for Captioning and Interpreting Memes

EunJeong Hwang, Vered Shwartz

Published: 07 Oct 2023, Last Modified: 01 Dec 2023EMNLP 2023 MainEveryoneRevisionsBibTeX

Submission Type: Regular Long Paper

Submission Track: Language Grounding to Vision, Robotics and Beyond

Submission Track 2: Speech and Multimodality

Keywords: meme, captioning, image, large multimodal model

Abstract: Memes are a widely popular tool for web users to express their thoughts using visual metaphors. Understanding memes requires recognizing and interpreting visual metaphors with respect to the text inside or around the meme, often while employing background knowledge and reasoning abilities. We present the task of meme captioning and release a new dataset, MemeCap. Our dataset contains 6.3K memes along with the title of the post containing the meme, the meme captions, the literal image caption, and the visual metaphors. Despite the recent success of vision and language (VL) models on tasks such as image captioning and visual question answering, our extensive experiments using state-of-the-art VL models show that they still struggle with visual metaphors, and perform substantially worse than humans.

Submission Number: 13

Loading