In-Image Neural Machine Translation with Segmented Pixel Sequence-to-Sequence Model

Yanzhi Tian; Xiang Li; Zeming Liu; Yuhang Guo; Bin Wang

In-Image Neural Machine Translation with Segmented Pixel Sequence-to-Sequence Model

Yanzhi Tian, Xiang Li, Zeming Liu, Yuhang Guo, Bin Wang

Published: 07 Oct 2023, Last Modified: 01 Dec 2023EMNLP 2023 FindingsEveryoneRevisionsBibTeX

Submission Type: Regular Long Paper

Submission Track: Machine Translation

Keywords: In-Image Machine Translation, Neural Machine Translation

TL;DR: We propose an end-to-end model for In-Image Machine Translation by encoding images to segmented pixel sequences.

Abstract: In-Image Machine Translation (IIMT) aims to convert images containing texts from one language to another. Traditional approaches for this task are cascade methods, which utilize optical character recognition (OCR) followed by neural machine translation (NMT) and text rendering. However, the cascade methods suffer from compounding errors of OCR and NMT, leading to a decrease in translation quality. In this paper, we propose an end-to-end model instead of the OCR, NMT and text rendering pipeline. Our neural architecture adopts encoder-decoder paradigm with segmented pixel sequences as inputs and outputs. Through end-to-end training, our model yields improvements across various dimensions, (i) it achieves higher translation quality by avoiding error propagation, (ii) it demonstrates robustness for out domain data, and (iii) it displays insensitivity to incomplete words. To validate the effectiveness of our method and support for future research, we construct our dataset containing 4M pairs of De-En images and train our end-to-end model. The experimental results show that our approach outperforms both cascade method and current end-to-end model.

Submission Number: 3530

Loading