Generalizing MLPs With Dropouts, Batch Normalization, and Skip Connections

Taewoon Kim

Generalizing MLPs With Dropouts, Batch Normalization, and Skip Connections

Taewoon Kim

Published: 28 Jan 2022, Last Modified: 22 Oct 2023ICLR 2022 SubmittedReaders: Everyone

Keywords: MLP, batch normalization, dropout, residual connections, Bayesian inference

Abstract: A multilayer perceptron (MLP) is typically made of multiple fully connected layers with nonlinear activation functions. There have been several approaches to make them better (e.g., faster convergence, better convergence limit, etc.). But the researches lack structured ways to test them. We test different MLP architectures by carrying out the experiments on the age and gender datasets. We empirically show that by whitening inputs before every linear layer and adding skip connections, our proposed MLP architecture can result in better performance. Since the whitening process includes dropouts, it can also be used to approximate Bayesian inference. We have open sourced our code, and released models and docker images at https://github.com/anonymous.

One-sentence Summary: By whitening inputs before every linear layer and adding skip connections, our proposed MLP architecture can result in better performance and capture uncertainty.

Community Implementations: [![CatalyzeX](/images/catalyzex_icon.svg) 2 code implementations](https://www.catalyzex.com/paper/arxiv:2108.08186/code)

5 Replies

Loading