arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18607cs.CV

VA-Judger:用于联合视频-音频生成的人类偏好反馈奖励建模

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

发表机构上海创新研究院 · 复旦大学 · 音网智能科技有限公司
查看机构详情
  • Shanghai Innovation Institution(上海创新研究院)
  • Fudan University(复旦大学)
  • Yinwang Intelligent Technology Co., Ltd(音网智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Hang Xu, Kaihang Pan, Yu-Gang Jiang, Zuxuan Wu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对联合视频-音频生成的奖励信号问题,构建了VAPref-10K数据集与VA-Judger-Bench基准,提出思维链全奖励模型VA-Judger,其在预测人类偏好及提升生成质量上均优于基线方法。

中文摘要 AI 辅助

使用强化学习对联合视频-音频生成模型进行后训练需要奖励信号。现有方法通过结合音频质量、视觉保真度、同步性等单个质量维度的指标来构建该奖励,但这些指标分别评估感知维度,无法捕捉文本提示、视频和音频之间决定人类偏好的整体语义与时间连贯性。针对这些指标优化模型会导致奖励黑客行为,生成的视频-音频内容在这些指标上得分高,但对人类观众而言不连贯或不忠实。为解决此问题,我们首先构建了用于联合视频-音频生成的大规模人类偏好数据集VAPref-10K,包含9K个提示和来自开源生成模型的10.3K个细粒度配对比较;还引入了VA-Judger-Bench基准,包含域内和域外模型比较,用于评估奖励模型是否真正与人类偏好对齐。我们进一步提出VA-Judger,一种用于联合视频-音频生成的思维链全奖励模型:首先从具有明显质量差距的配对中学习以建立结构化输出和粗粒度偏好判别,然后通过经人类注释验证的拒绝采样,为更难的近质量比较提炼可靠的偏好解释,最后执行维度分解的强化学习,将人类反馈分解为单个质量维度,以获得比单一二元偏好标签更密集的奖励信号。实验表明,VA-Judger在域内和域外评估中预测人类偏好的性能均优于指标基线,使用其人类对齐的奖励对音频-视频生成模型进行后训练也能在生成质量上取得显著提升。

英文摘要

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.

补充信息

↑