对齐人类感知:用于视频生成的校准分布奖励学习
Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation
- City University of Hong Kong(香港城市大学)
- ShanghaiTech University(上海科技大学)
- University of California, Berkeley(加利福尼亚大学伯克利分校)
- The Ohio State University(俄亥俄州立大学)
- University of Glasgow(格拉斯哥大学)
- Columbia University(哥伦比亚大学)
- University of Arizona(亚利桑那大学)
- GMI Cloud
- National University of Singapore(新加坡国立大学)
- DeciLix Lab
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对视频生成中奖励信号不可靠、偏好维度权衡丢失、KL散度难捕捉偏好全局结构的问题,提出统一偏好感知学习框架,通过精英过滤、分布建模与Wasserstein对齐改进,提升了奖励可靠性与视频感知一致性。
AI中文摘要:
视频生成是AI驱动内容创作的核心,将生成视频与人类偏好对齐是评估生成质量的关键标准。尽管视觉质量取得了显著进展,但仍存在三大关键挑战:其一,奖励信号的可靠性受限于人类偏好数据的质量,这类数据常受主观噪声与偏差影响;其二,标准标量奖励模型将多维度人类偏好压缩为单一值,导致跨多个偏好维度的动态权衡信息丢失;其三,在策略优化中,广泛采用的KL散度主要施加局部约束,可能无法捕捉人类偏好的全局结构。为应对这些挑战,本文提出一种用于视频生成的统一偏好感知学习框架:首先引入精英引导过滤以校准偏好数据,为奖励模型训练构建可靠监督;接着将视频质量建模为多维奖励分布,以捕捉人类偏好固有的不确定性,并使用Wasserstein距离使学习到的奖励分布与经验人类偏好分布对齐;最后将基于Wasserstein的分布对齐引入GRPO,指导策略优化更好地匹配人类对视频偏好的全局结构。在奖励建模与视频生成上的实验表明,所提方法提升了奖励信号的可靠性与生成视频的感知一致性,代码可在指定URL获取。
英文摘要:
Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating generation quality. Despite significant progress in visual quality, three key challenges remain. First, the reliability of reward signals is constrained by the quality of human preference data, which is often affected by subjective noise and bias. Second, standard scalar reward models collapse multi-aspect human preferences into a single value, leading to the loss of dynamic trade-offs across multiple preference dimensions. Third, in policy optimization, the widely adopted KL divergence imposes primarily local constraints and may fail to capture the global structure of human preferences. To address these challenges, we propose a unified preference-aware learning framework for video generation. First, we introduce elite-guided filtering to calibrate preference data and construct reliable supervision for reward model training. We then model video quality as a multidimensional reward distribution to capture the uncertainty inherent in human preferences, and use the Wasserstein distance to align the learned reward distribution with the empirical human preference distribution. Finally, we introduce Wasserstein-based distributional alignment into GRPO, guiding policy optimization to better match the global structure of human preferences over videos. Experiments on reward modeling and video generation demonstrate that our approach improves the reliability of reward signals and the perceptual consistency of generated videos. Our code is available at https://github.com/alignhs26/ahs.