Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets

Published: 21 Feb 2025, Last Modified: 21 Feb 2025RLGMSD 2024 PosterEveryoneRevisionsBibTeXCC BY 4.0
Keywords: Contrastive Learning, Tabular Data, Privacy Metrics
TL;DR: This work presents a contrastive learning-based method to improve privacy risk assessment in synthetic tabular data, addressing GDPR-defined "singling out" risks with efficient and effective metrics.
Abstract: Synthetic data has garnered attention as a Privacy Enhancing Technology in sectors such as healthcare and finance. When using synthetic data in practical applications, it is important to provide protection guarantees. We introduce a contrastive method that improves privacy assessment of synthetic datasets by embedding the data in a more representative space. This overcomes obstacles surrounding the multitude of data types and attributes. It also makes the use of intuitive distance metrics possible for similarity measurements and as an attack vector. Our results show that relatively efficient, easy to implement privacy metrics can perform equally well as more advanced metrics explicitly modeling conditions for privacy referred to by the GDPR.
Submission Number: 4
Loading

OpenReview is a long-term project to advance science through improved peer review with legal nonprofit status. We gratefully acknowledge the support of the OpenReview Sponsors. © 2025 OpenReview