Egocentric Cross-Embodiment Video Editing via Dual Contrastive Representation Learning

Zhiyuan Li; Wenyan Yang; Wenshuai Zhao; Yue Ma; Yuanpeng Tu; Pekka Marttinen; Joni Pajarinen

Egocentric Cross-Embodiment Video Editing via Dual Contrastive Representation Learning

Zhiyuan Li, Wenyan Yang, Wenshuai Zhao, Yue Ma, Yuanpeng Tu, Pekka Marttinen, Joni Pajarinen

17 Sept 2025 (modified: 11 Feb 2026)Submitted to ICLR 2026EveryoneRevisionsBibTeXCC BY 4.0

Keywords: video generation, diffusion transformer, egocentric videos, cross-embodiment gap, contrastive learning

TL;DR: We achieve egocentric cross-embodiment video editing by using a dual contrastive objective to disentangle a human demonstration and generate a coherent, robot-centric video.

Abstract: Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled representations, where task-relevant information is coupled with human-specific kinematics, limiting their adaptability. We propose a generative framework for cross-embodiment video editing that directly addresses this by learning explicitly disentangled task and embodiment representations. Our method factorizes a demonstration video into two orthogonal latent spaces by enforcing a dual contrastive objective: it minimizes mutual information between the spaces to ensure independence while maximizing intra-space consistency to create stable representations. A parameter-efficient adapter injects these latent codes into a frozen video diffusion model, enabling the synthesis of a coherent robot execution video from a single human demonstration, without requiring paired cross-embodiment data. Experiments show our approach generates temporally consistent and morphologically accurate robot demonstrations, offering a scalable solution to leverage internet-scale human video for robot learning.

Supplementary Material: zip

Primary Area: generative models

Submission Number: 9805

Loading