Foundation Models for Clinical Records at Health System Scale

Haresh Rengaraj Rajamohan; Xiang Gao; Weicheng Zhu; Shih-Lun Huang; Long Chen; Kyunghyun Cho; Cem M Deniz; Narges Razavian

Foundation Models for Clinical Records at Health System Scale

Haresh Rengaraj Rajamohan, Xiang Gao, Weicheng Zhu, Shih-Lun Huang, Long Chen, Kyunghyun Cho, Cem M Deniz, Narges Razavian

Published: 09 Jun 2025, Last Modified: 01 Jul 2025FMSD @ ICML 2025EveryoneRevisionsBibTeXCC BY 4.0

Keywords: foundation models, electronic health records, generative pretraining, zero-shot prediction

TL;DR: We introduce a generative pretraining framework for EHR foundation models that predicts next-visit clinical events and incorporates a repeat-aware regularization scheme to improve forecasting of new disease onsets in a zero-shot setting.

Abstract: Large-scale pretraining has transformed modeling of language and other data types, but its potential remains underexplored in healthcare with structured electronic health records (EHRs). We present a novel generative pretraining strategy for sequential EHR data using next-visit event prediction. Our model learns to autoregressively generate various tokenized clinical events for the next visit based on patient history and inherently handles the joint prediction of heterogeneous data types. Additionally, we introduce regularization on predicting repeated events and highlight a key pitfall in EHR-based foundation model evaluations: repeated event tokens can inflate performance metrics when new onsets are not distinguished from subsequent occurrences. Our model is evaluated via zero-shot prediction for forecasting dementia and knee osteoarthritis incidence within 2 and 5 years, and the model performance rivals a fully fine-tuned masked pretrained Transformer baseline, demonstrating that our approach captures complex clinical dependencies without requiring costly task-specific fine-tuning.

Submission Number: 59

Loading