Continual Transformers: Redundancy-Free Attention for Online Inference

Lukas Hedegaard; Arian Bakhtiarnia; Alexandros Iosifidis

Continual Transformers: Redundancy-Free Attention for Online Inference

Lukas Hedegaard, Arian Bakhtiarnia, Alexandros Iosifidis

Published: 01 Feb 2023, Last Modified: 22 Jun 2025ICLR 2023 posterReaders: Everyone

Keywords: Transformer, Continual Inference Networks, Online inference, Stream processing, Acceleration, Online Action Detection, Audio classification

TL;DR: A Transformer Decorder acceleration for online stream processing validated with experiments in Online Action Detection and Audio Classification.

Abstract: Transformers in their common form are inherently limited to operate on whole token sequences rather than on one token at a time. Consequently, their use during online inference on time-series data entails considerable redundancy due to the overlap in successive token sequences. In this work, we propose novel formulations of the Scaled Dot-Product Attention, which enable Transformers to perform efficient online token-by-token inference on a continual input stream. Importantly, our modifications are purely to the order of computations, while the outputs and learned weights are identical to those of the original Transformer Encoder. We validate our Continual Transformer Encoder with experiments on the THUMOS14, TVSeries and GTZAN datasets with remarkable results: Our Continual one- and two-block architectures reduce the floating point operations per prediction by up to 63x and 2.6x, respectively, while retaining predictive performance.

Anonymous Url: I certify that there is no URL (e.g., github page) that could be used to find authors’ identity.

No Acknowledgement Section: I certify that there is no acknowledgement section in this submission for double blind review.

Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics

Submission Guidelines: Yes

Please Choose The Closest Area That Your Submission Falls Into: Deep Learning and representational learning

Community Implementations: [![CatalyzeX](/images/catalyzex_icon.svg) 1 code implementation](https://www.catalyzex.com/paper/continual-transformers-redundancy-free/code)

9 Replies

Loading