VideoScore2: Think before You Score in Generative Video Evaluation

Xuan He; Dongfu Jiang; Ping Nie; Minghao Liu; Zhengxuan Jiang; Mingyi Su; Wentao Ma; Junru Lin; Chun Ye; Yi Lu; Keming Wu; Benjamin Schneider; Quy Duc Do; Zhuofeng Li; Yiming Jia; Yuxuan Zhang; Guo Cheng; Haozhe Wang; Wangchunshu Zhou; Qunshu Lin; Yuanxing Zhang; Ge Zhang; Wenhao Huang; Wenhu Chen

VideoScore2: Think before You Score in Generative Video Evaluation

18 Sept 2025 (modified: 13 Nov 2025)ICLR 2026 Conference Withdrawn SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Keywords: Video Quality Evaluation, Reward Modeling, Multimodal Large Language Models, Reinforcement Learning

TL;DR: We present VideoScore2, a model able to evaluate generative videos with long-CoT thinking process.

Abstract: Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present **VideoScore2**, a *multi-dimensional*, *interpretable*, and *human-aligned* framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset **VideoFeedback2** containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark VideoScore-Bench-v2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling.

Primary Area: applications to computer vision, audio, language, and other modalities

Submission Number: 13258

Loading

VideoScore2: Think before You Score in Generative Video Evaluation

Xuan He, Dongfu Jiang, Ping Nie, Minghao Liu, Zhengxuan Jiang, Mingyi Su, Wentao Ma, Junru Lin, Chun Ye, Yi Lu, Keming Wu, Benjamin Schneider, Quy Duc Do, Zhuofeng Li, Yiming Jia, Yuxuan Zhang, Guo Cheng, Haozhe Wang, Wangchunshu Zhou, Qunshu Lin et al. (4 additional authors not shown)