arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当物理偏好遇到语义约束:用于文本到视频生成的物理和语义直接偏好优化

When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

Siwei Meng, Yawei Luo, Shu Zhang, Ping Liu

arXiv 2607.16947首次发表:更新:

发表机构

University of Nevada, Reno; Zhejiang University(内华达大学雷诺分校; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

文本到视频生成模型存在物理合理性与语义一致性的矛盾,提出物理和语义直接偏好优化方法(PSDPO),通过调节偏好对贡献,限制语义漂移,实现更可靠平衡,实验验证其在提高物理合理性和保持语义一致性上优于基线和现有方法。

AI 中文摘要

文本到视频(T2V)生成模型已实现强大的视觉逼真度,但提高物理合理性可能会以与输入文本的语义一致性为代价。这种矛盾的产生是因为物理偏好通常通过比较两个视频之间的动态来确定,而不考虑任何一个视频是否忠实地描绘了提示指定的场景,这使得物理语义冲突在这种监督范式下成为一种系统趋势。我们将此挑战表述为一个约束偏好优化问题,并提出物理和语义直接偏好优化(PSDPO),它基于每个偏好对的物理和语义信号之间的一致性来调节其贡献。梯度级分析表明,PSDPO将冲突对的语义漂移限制在可控的残差范围内,并进一步推动了一种可证明减少累积漂移的分阶段优化协议。实验表明,PSDPO在VideoPhy-2上比基线提高了高达2倍的物理合理性,同时在VBench上保持了强大的语义一致性,比现有的基于偏好的方法实现了更可靠的平衡。

英文摘要

Text-to-video (T2V) generation models have achieved strong visual realism, but improving physical plausibility can come at the cost of semantic consistency with the input text. This tension arises because physical preference is typically determined by comparing dynamics between two videos, without accounting for whether either video faithfully depicts the scene specified by the prompt, making physical-semantic conflict a systematic tendency under this supervision paradigm. We formulate this challenge as a constrained preference optimization problem and propose Physical and Semantic Direct Preference Optimization (PSDPO), which modulates each preference pair's contribution based on the agreement between its physical and semantic signals. A gradient-level analysis shows that PSDPO bounds the semantic drift from conflicting pairs to a controllable residual, and further motivates a staged optimization protocol that provably reduces cumulative drift. The resulting method operates entirely within the standard DPO framework, requiring no auxiliary models or additional loss terms. Experiments show that PSDPO improves physical plausibility by up to $2\times$ over the baseline on VideoPhy-2, while maintaining strong semantic consistency on VBench, achieving a more reliable balance than existing preference-based methods.

CommentsThis work is accepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑