Tradeoffs in Data Augmentation: An Empirical Study

Raphael Gontijo-Lopes; Sylvia Smullin; Ekin Dogus Cubuk; Ethan Dyer

Tradeoffs in Data Augmentation: An Empirical Study

Raphael Gontijo-Lopes, Sylvia Smullin, Ekin Dogus Cubuk, Ethan Dyer

Published: 12 Jan 2021, Last Modified: 05 May 2023ICLR 2021 PosterReaders: Everyone

Keywords: Generalization, Interpretability, Understanding Data Augmentation

Abstract: Though data augmentation has become a standard component of deep neural network training, the underlying mechanism behind the effectiveness of these techniques remains poorly understood. In practice, augmentation policies are often chosen using heuristics of distribution shift or augmentation diversity. Inspired by these, we conduct an empirical study to quantify how data augmentation improves model generalization. We introduce two interpretable and easy-to-compute measures: Affinity and Diversity. We find that augmentation performance is predicted not by either of these alone but by jointly optimizing the two.

One-sentence Summary: We quantify mechanisms of how data augmentation works with two metrics we introduce: Affinity and Diversity.

Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics

Supplementary Material: zip

17 Replies

Loading