发表机构
Zhejiang University; The Hong Kong Polytechnic University; IROOTECH TECHNOLOGY(浙江大学; 香港理工大学; 易路泰克科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视听生成模型跨模态同步评估难题,提出框架解决人类与自动化评估指标的矛盾。通过引入数据集、适配模型及提出优化方法,实现先进的人类偏好对齐,建立标准化基准推动评估从信号相关性到因果关系发展。
AI 中文摘要
随着视听生成模型演变成世界模拟器,跨模态同步成为评估生成内容中世界动态和因果关系一致性的关键代理。然而,现有评估指标假定结构正确,将同步简化为单纯的时间对齐。因此,它们在生成输出上失败,特别是当出现结构幻觉和不对称跨模态关系时,目前这需要专家人工标注来评估同步。这种依赖引入了一个关键悖论:人类评估者依赖相对的、依赖参考的比较,而自动化指标需要无参考的、绝对标量。我们通过提出一个框架来解决这个悖论,该框架将相对人类感知提炼为一个连续的、全局一致的指标。首先,我们引入SynthSync,一个通过成对人工标注排序的生成失败数据集。其次,我们适配配备连续潜在投影的Omni-LLM,将相对人类排名转换为连续绝对值。第三,我们提出实值组相对策略优化($\mathbb{R}$-GRPO),通过逐列表评分分布内化同步的全局因果结构。从经验上看,我们的指标实现了最先进的人类偏好对齐。我们利用这个估计器建立一个标准化基准,将AV-Gen评估从低级信号相关性推进到视觉基础的因果关系。
英文摘要
As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality in generated content. However, existing evaluation metrics presume structural correctness, reducing synchronization to mere temporal alignment. Consequently, they fail on generative outputs, especially when exhibiting structural hallucinations and asymmetric cross-modal relations, which currently \textbf{mandate expert human annotation to assess synchronization.} This dependency introduces a critical paradox: \emph{human evaluators rely on relative, reference-dependent comparisons, whereas automated metrics require reference-free, absolute scalars.} We resolve this paradox by proposing a framework that distills relative human perception into a continuous, globally consistent metric. First, we introduce SynthSync, a dataset of generative failures ranked via pairwise human annotations. Second, we adapt the Omni-LLM equipped with a continuous latent projection to translate relative human rankings into continuous absolute values. Third, we propose Real-Valued Group Relative Policy Optimization ($\mathbb{R}$-GRPO) to internalize the global causal structure of synchronization via listwise score distributions. Empirically, our metric achieves state-of-the-art human preference alignment. We leverage this estimator to establish a standardized benchmark, advancing AV-Gen assessment from low-level signal correlation to visually grounded causality.
CommentsAccepted to ECCV 2026