Model-Based Data-Centric AI: Bridging the Divide Between Academic Ideals and Industrial Pragmatism

Model-Based Data-Centric AI: Bridging the Divide Between Academic Ideals and Industrial Pragmatism

ICLR 2024 Workshop DMLR Submission35 Authors

Published: 04 Mar 2024, Last Modified: 02 May 2024DMLR @ ICLR 2024EveryoneRevisionsBibTeXCC BY 4.0

Keywords: Data-Centric AI, Model-Agnostic AI, Model-Based Data-Centric AI

TL;DR: This paper examines the differences between Data-Centric AI and Model-Agnostic AI, highlighting the challenges and proposing a new approach for integrating model considerations into data optimization to align academic and industrial standards

Abstract: This paper delves into the contrasting roles of data within academic and industrial spheres, highlighting the divergence between Data-Centric AI and Model-Agnostic AI approaches. We argue that while Data-Centric AI focuses on the primacy of high-quality data for model performance, Model-Agnostic AI prioritizes algorithmic flexibility, often at the expense of data quality considerations. This distinction reveals that academic standards for data quality frequently do not meet the rigorous demands of industrial applications, leading to potential pitfalls in deploying academic models in real-world settings. Through a comprehensive analysis, we address these disparities, presenting both the challenges they pose and strategies for bridging the gap. Furthermore, we propose a novel paradigm: Model-Based Data-Centric AI, which aims to reconcile these differences by integrating model considerations into data optimization processes. This approach underscores the necessity for evolving data requirements that are sensitive to the nuances of both academic research and industrial deployment. By exploring these discrepancies, we aim to foster a more nuanced understanding of data's role in AI development and encourage a convergence of academic and industrial standards to enhance AI's real-world applicability.

Primary Subject Area: Other

Paper Type: Research paper: up to 8 pages

DMLR For Good Track: Participate in DMLR for Good Track

Participation Mode: Virtual

Confirmation: I have read and agree with the workshop's policy on behalf of myself and my co-authors.

Submission Number: 35

Loading