MuseTalk: Real-Time High Quality Lip Synchronization with Latent Space Inpainting

Yue Zhang; Minhao LIU; Zhaokang Chen; Bin Wu; Yubin zeng; Chao Zhan; Yingjie He; JUNXIN HUANG; Wenjiang Zhou

MuseTalk: Real-Time High Quality Lip Synchronization with Latent Space Inpainting

Yue Zhang, Minhao LIU, Zhaokang Chen, Bin Wu, Yubin zeng, Chao Zhan, Yingjie He, JUNXIN HUANG, Wenjiang Zhou

27 Sept 2024 (modified: 13 Nov 2024)ICLR 2025 Conference Withdrawn SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Keywords: talking face, face visual dubbing, generative models, multimodality, AI-generated content

TL;DR: We propose MuseTalk, which generate lip-sync targets in latent space encoded by a Variational Autoencoder. Generating in latent space allows MuseTalk to operate in real-time and to generate high-resolution face images.

Abstract: Achieving high-resolution, identity consistency, and accurate lip-speech synchronization in face visual dubbing presents significant challenges, particularly for real-time applications like live video streaming. We propose MuseTalk, which generates lip-sync targets in a latent space encoded by a Variational Autoencoder, enabling high-fidelity talking face video generation with efficient inference. Specifically, we project the occluded lower half of the face image and itself as an reference into a low-dimensional latent space and use a multi-scale U-Net to fuse audio and visual features at various levels. We further propose a novel sampling strategy during training, which selects reference images with head poses closely matching the target, allowing the model to focus on precise lip movement by filtering out redundant information. Additionally, we analyze the mechanism of lip-sync loss and reveal its relationship with input information volume. Extensive experiments show that MuseTalk consistently outperforms recent state-of-the-art methods in visual fidelity and achieves comparable lip-sync accuracy. As MuseTalk supports the online generation of face at 256x256 at more than 30 FPS with negligible starting latency, it paves the way for real-time applications. The codes and models will be made publicly available upon acceptance.

Supplementary Material: zip

Primary Area: generative models

Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics.

Submission Guidelines: I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide.

Reciprocal Reviewing: I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6.

Anonymous Url: I certify that there is no URL (e.g., github page) that could be used to find authors’ identity.

No Acknowledgement Section: I certify that there is no acknowledgement section in this submission for double blind review.

Submission Number: 9580

Loading