arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhyProbe:重新思考生成视频中的物理一致性评估

PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos

Max Ku, Jiaojiao Fan, Zekun Hao, Francesco Ferroni, Heng Wang, Wenhu Chen, Ming-Yu Liu, Prithvijit Chattopadhyay

arXiv 2609.38377首次发表:更新:

发表机构

NVIDIA; University of Waterloo; Vector Institute(英伟达; 滑铁卢大学; 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhyProbe利用冻结的预训练时空编码器和轻量级评分头,通过统一训练目标评估生成视频的物理一致性,在多个基准上超越现有方法,并与人类判断高度相关。

AI 中文摘要

评估生成视频的物理一致性仍然是一个基本挑战。现有方法依赖于现成的视觉-语言模型,这些模型往往对物理动态缺乏全局视野,或者依赖于在人工标注上微调的评估器,这些评估器会过度拟合数据集特定的线索,并且难以泛化。一个关键挑战是,现有的监督源要么提供相对排序,要么提供绝对分数,但不能在多种设置下同时可靠且一致地提供两者。为此,我们引入了PhyProbe,一种评估器,它从冻结的预训练时空编码器中提取特征,并通过轻量级评分头将其映射到标量物理一致性违规分数。PhyProbe通过统一目标进行训练,该目标结合了成对排序、对噪声标量注释的回归,以及基于锚点的校准,覆盖了一组精心挑选的异构监督源。实验表明,PhyProbe在大多数成对基准上优于先前方法,这些基准涵盖真实-生成和生成-生成对,且对应关系各异,在无对应关系和生成-生成设置中增益最大,而现有微调评估器在这些设置中性能急剧下降。PhyProbe与人类判断实现了强相关性,排名基和线性指标之间高度一致,表明分数既排序良好,又锚定在稳定的[0, 1]尺度上。此外,尽管PhyProbe是在指示物理一致性的监督下训练的,没有显式的通用偏好标签,它在人类偏好基准上也表现出竞争力:这与物理违规与更广泛的质量退化纠缠在一起的观察结果一致。

英文摘要

Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe outperforms prior methods on most pairwise benchmarks spanning real-generated and generated-generated pairs under varying correspondence, with the largest gains in no-correspondence and generated-generated settings where existing fine-tuned evaluators degrade sharply. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating that scores are both well ordered and anchored to a stable [0, 1] scale. Further, despite being trained on supervision indicative of physical consistency, without explicit general-preference labels, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.

CommentsNeurIPS 2026 poster

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑