Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?

Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?

ACL ARR 2025 May Submission6821 Authors

20 May 2025 (modified: 03 Jul 2025)ACL ARR 2025 May SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Abstract: As large language models (LLMs) continue to improve, their evaluation increasingly centers on complex, high-level tasks, often at the expense of systematically assessing fundamental capabilities. To address this gap, recent work proposed LMentry, a compact benchmark comprising tasks that are trivial for humans but remain surprisingly difficult for LLMs. However, LMentry is limited to English, leaving its insights linguistically narrow. In this paper, we present Multi-LMentry, a ground-up recreation of LMentry that enables systematic evaluation of LLMs on basic reasoning and understanding tasks across nine diverse languages. Multi-LMentry includes English and expands to Catalan, German, Spanish, Basque, Galician, Korean, Italian, and Brazilian Portuguese, emphasizing the importance of cross-lingual and low-resource settings. To validate that Multi-LMentry is still trivial for humans, we demonstrate that L2 speakers with only elementary proficiency achieve near-perfect scores in a low-resource language, namely, Basque. Through extensive experiments, we reveal that state-of-the-art open-weight multilingual LLMs still fall short of human performance on elementary tasks in many languages. Our results expose new failure modes that remain hidden in monolingual evaluation, underscoring the need for rigorous, language-diverse ``unit tests'' of core model abilities.

Paper Type: Long

Research Area: Resources and Evaluation

Research Area Keywords: benchmarking, language resources, multilingual corpora, datasets for low resource languages

Contribution Types: Approaches to low-resource settings, Data resources

Languages Studied: English, Italian, Spanish, German, Catalan, Korean, Basque, Galician, Portuguese (Brazilian),

Submission Number: 6821

Loading