Unsupervised Feature Learning for Audio Analysis

Matthias Meyer; Jan Beutel; Lothar Thiele

Unsupervised Feature Learning for Audio Analysis

Matthias Meyer, Jan Beutel, Lothar Thiele

07 Jun 2026 (modified: 17 Feb 2017)ICLR 2017Readers: Everyone

Abstract: Identifying acoustic events from a continuously streaming audio source is of interest for many applications including environmental monitoring for basic research. In this scenario neither different event classes are known nor what distinguishes one class from another. Therefore, an unsupervised feature learning method for exploration of audio data is presented in this paper. It incorporates the two following novel contributions: First, an audio frame predictor based on a Convolutional LSTM autoencoder is demonstrated, which is used for unsupervised feature extraction. Second, a training method for autoencoders is presented, which leads to distinct features by amplifying event similarities. In comparison to standard approaches, the features extracted from the audio frame predictor trained with the novel approach show 13 % better results when used with a classifier and 36 % better results when used for clustering.

TL;DR: Novel method to train autoencoders by introducing pairwise loss. Use case: Unsupervised audio analysis with an audio frame predictor using Convolutional LSTM layers

Conflicts: matthias.meyer@tik.ee.ethz.ch, beutel@tik.ee.ethz.ch, thiele@tik.ee.ethz.ch

5 Replies

Loading