Conditional Synthetic Data Generation for Robust Machine Learning Applications with Limited Pandemic Data

Hari Prasanna Das; Ryan Tran; Japjot Singh; Xiangyu Yue; Geoffrey H. Tison; Alberto Sangiovanni-Vincentelli; Costas Spanos

Conditional Synthetic Data Generation for Robust Machine Learning Applications with Limited Pandemic Data

Hari Prasanna Das, Ryan Tran, Japjot Singh, Xiangyu Yue, Geoffrey H. Tison, Alberto Sangiovanni-Vincentelli, Costas Spanos

Published: 30 Jul 2022, Last Modified: 17 May 2023KDD 2022 Workshop epiDAMIK OralReaders: Everyone

Keywords: COVID, Pandemic, Conditional Synthetic Data Generation

TL;DR: To tackle the challenges of limited pandemic data, and label scarcity in the available data, we propose generating conditional synthetic data using a hybrid model consisting of a conditional generative flow and a classifier.

Abstract: $\textbf{Background:}$ At the onset of a pandemic, such as COVID-19, data with proper labeling/attributes corresponding to the new disease might be unavailable or sparse. Machine Learning (ML) models trained with the available data, which is limited in quantity and poor in diversity, will often be biased and inaccurate. At the same time, ML algorithms designed to fight pandemics must have good performance and be developed in a time-sensitive manner. To tackle the challenges of limited data, and label scarcity in the available data, we propose generating conditional synthetic data, to be used alongside real data for developing robust ML models. $\textbf{Methods:}$ We present a hybrid model consisting of a conditional generative flow and a classifier for conditional synthetic data generation. $\textbf{Results:}$ We performed conditional synthetic generation for chest computed tomography (CT) scans corresponding to normal, COVID-19, and pneumonia afflicted patients. We show that our method significantly outperforms existing models both on qualitative and quantitative performance. As an example of downstream use of synthetic data, we show improvement in COVID-19 detection from CT scans with conditional synthetic data augmentation.

3 Replies

Loading