Latent Space Structuring for Conditional Tabular Data Generation on Imbalanced Datasets

Latent Space Structuring for Conditional Tabular Data Generation on Imbalanced Datasets

ICLR 2026 Conference Submission14087 Authors

18 Sept 2025 (modified: 08 Oct 2025)ICLR 2026 Conference SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Keywords: synthetic tabular data, class imbalance, conditional generation, transformer-based models

TL;DR: Generating useful, faithful and diverse data with a focus on minority samples in imbalanced settings.

Abstract: Generating synthetic tabular data under severe class imbalance is essential for domains where rare but high-impact events drive decision-making. Yet most generative models either overlook minority groups or fail to produce samples that are useful for downstream learning. We introduce CTTVAE, a Conditional Transformer-based Tabular Variational Autoencoder equipped with two complementary mechanisms: (i) a class-aware triplet margin loss that restructures the latent space for sharper intra-class compactness and inter-class separation, and (ii) a training-by-sampling strategy that adaptively increases exposure to underrepresented groups. Together, these components form CTTVAE+TBS, a framework that consistently yields more representative and utility-aligned samples without destabilizing training. Across six real-world benchmarks, CTTVAE+TBS achieves the strongest downstream utility on minority classes, often surpassing models trained on the original imbalanced data while maintaining competitive fidelity and privacy. Ablation studies further confirm that both latent structuring and targeted sampling contribute to these gains. By explicitly prioritizing downstream performance in rare categories, CTTVAE+TBS provides a robust and interpretable solution for conditional tabular data generation, with direct applicability to industries like healthcare, fraud detection, and predictive maintenance where even small gains on minority cases can be critical.

Supplementary Material: zip

Primary Area: generative models

Submission Number: 14087

Loading