Aloe-Vision: Robust Vision-Language Models for Healthcare

Jaume Guasch-Martí; Enrique Lopez-Cuena; Martín Suárez-Fernández; Jordi Bayarri-Planas; Anna Arias-Duart; Dario Garcia-Gasulla

Aloe-Vision: Robust Vision-Language Models for Healthcare

Jaume Guasch-Martí, Enrique Lopez-Cuena, Martín Suárez-Fernández, Jordi Bayarri-Planas, Anna Arias-Duart, Dario Garcia-Gasulla

Published: 14 Feb 2026, Last Modified: 14 Feb 2026MIDL 2026 PosterEveryoneRevisionsBibTeXCC BY 4.0

Keywords: LVLMs, Healthcare, Adversarial evaluation

Abstract: Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the scarcity of high-quality medical multimodal data, concerns about robustness in safety-critical settings, and the narrow and potentially contaminated evaluation benchmarks that limit reliable assessment. To address these issues, the field requires state-of-the-art solutions to be fully open and reproducible systems in which all components can be inspected, evaluated, and improved. This work introduces the Aloe-Vision family of medical LVLMs, openly released with full weights, training recipes and data, in two scales (7B and 72B). The models are trained on Aloe-Vision-Data, our quality-filtered mixture which integrates both medical and general domains across multimodal and text-only sources, designed for direct use in model fine-tuning. Through comprehensive benchmarking, we demonstrate that balanced training mixtures produce robust LVLMs which yield significant gains over the baseline models without compromising general capabilities, achieving competitive performance with state-of-the-art alternatives. We also introduce CareQA-Vision, a carefully curated vision benchmark derived from MIR and EIR exams, the residency entrance exams for medical and nursing specialists in Spain, offering novel vision questions with minimal contamination. Finally, we show that current LVLMs remain vulnerable to adversarial and misleading inputs, underscoring reliability challenges in clinical contexts.

Primary Subject Area: Foundation Models

Secondary Subject Area: Safe and Trustworthy Learning-assisted Solutions for Medical Imaging

Registration Requirement: Yes

Read CFP & Author Instructions: Yes

Originality Policy: Yes

Single-blind & Not Under Review Elsewhere: Yes

LLM Policy: Yes

Submission Number: 206

Loading