Limited Linguistic Diversity in Embodied AI Datasets

Limited Linguistic Diversity in Embodied AI Datasets

ACL ARR 2026 January Submission9516 Authors

06 Jan 2026 (modified: 20 Mar 2026)ACL ARR 2026 January SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Keywords: vision-language-action models, dataset audit, linguistic diversity

Abstract: Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate these systems remain poorly documented. In this work, we present a systematic dataset audit of several widely used VLA corpora, aiming to characterize what kinds of instructions these datasets actually contain and how much linguistic variety they provide. We quantify instruction language along complementary dimensions—including lexical variety, duplication and overlap, semantic similarity, and syntactic complexity. Our analysis shows that many datasets rely on highly repetitive, template-like commands with limited structural variation, yielding a narrow distribution of instruction forms. We position these findings as descriptive documentation of the language signal available in current VLA training and evaluation data, intended to support more detailed dataset reporting, more principled dataset selection, and targeted curation or augmentation strategies that broaden language coverage.

Paper Type: Long

Research Area: Resources and Evaluation

Research Area Keywords: automatic evaluation of datasets

Contribution Types: Data analysis, Position papers

Languages Studied: English

Submission Number: 9516

Loading