发表机构
Cornell University; The University of Hong Kong; Johns Hopkins University(康奈尔大学; 香港大学; 约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频扩散模型中的潜在奖励黑客问题,提出CoRe协同进化奖励框架,通过动态重拟合奖励模型并锚定真实偏好,提升生成质量并避免质量崩溃。
AI 中文摘要
潜在奖励模型(LRMs)通过在潜在空间中直接对中间状态进行评分,实现了视频扩散模型的高效对齐。然而,我们发现针对固定潜在奖励进行优化会迅速导致潜在奖励黑客问题:预测奖励保持较高,而感知质量和运动质量却下降。我们的分析将分布逃逸确定为核心原因:在几百次更新内,生成器超出了奖励模型的训练支持范围,此时其评分不再反映视频质量。基于这一见解,我们引入了CoRe,一个协同进化奖励框架,将潜在空间对齐视为生成器与奖励模型之间的动态交互。CoRe不是针对静态代理进行优化,而是不断在生成器当前样本上重新拟合奖励模型,同时将其锚定到真实视频偏好上,从而使得生成器无法通过偏离数据来获得奖励。在Wan2.1-T2V-1.3B上,实验表明,CoRe在生成质量上始终优于预训练模型和先前的对齐方法,同时避免了固定奖励优化的质量崩溃问题。
英文摘要
Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.