Monotonic Multihead Attention

Xutai Ma; Juan Miguel Pino; James Cross; Liezl Puzon; Jiatao Gu

Monotonic Multihead Attention

Xutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon, Jiatao Gu

Published: 20 Dec 2019, Last Modified: 22 Jun 2025ICLR 2020 Conference Blind SubmissionReaders: Everyone

TL;DR: Make the transformer streamable with monotonic attention.

Abstract: Simultaneous machine translation models start generating a target sequence before they have encoded or read the source sequence. Recent approach for this task either apply a fixed policy on transformer, or a learnable monotonic attention on a weaker recurrent neural network based structure. In this paper, we propose a new attention mechanism, Monotonic Multihead Attention (MMA), which introduced the monotonic attention mechanism to multihead attention. We also introduced two novel interpretable approaches for latency control that are specifically designed for multiple attentions. We apply MMA to the simultaneous machine translation task and demonstrate better latency-quality tradeoffs compared to MILk, the previous state-of-the-art approach.

Keywords: Simultaneous Translation, Transformer, Monotonic Attention

Code: [![github](/images/github_icon.svg) pytorch/fairseq](https://github.com/pytorch/fairseq) + [![Papers with Code](/images/pwc_icon.svg) 2 community implementations](https://paperswithcode.com/paper/?openreview=Hyg96gBKPS)

Community Implementations: [![CatalyzeX](/images/catalyzex_icon.svg) 1 code implementation](https://www.catalyzex.com/paper/monotonic-multihead-attention/code)

Original Pdf: pdf

15 Replies

Loading