arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03952cs.CV

WorldReward:面向相机条件世界模型的奖励建模

WorldReward: Reward Modeling for Camera-Conditioned World Models

  • Fudan University(复旦大学)
  • Tencent Hunyuan(腾讯混元)
  • Shanghai Innovation Institute(上海创新研究院)
  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Y… 展开作者

Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang

AI总结:

该研究提出WorldReward,一种基于VLM的成对偏好奖励模型,解决相机条件世界模型现有奖励函数的缺陷,在WorldReward-Bench上超GPT-5.5,还提升了HY-WorldPlay 1.5的动作与视觉质量。

AI中文摘要:

相机条件世界模型生成交互式视频,要求指令动作能引发预期的场景变化,同时外观、几何结构和时间动态保持一致。现有奖励函数分别评估这些需求:基于几何的奖励估计轨迹执行情况,但无法判断执行动作的视觉质量;而基于图像的奖励测量帧质量,却无法捕捉动作执行或时间动态。我们假设视觉语言模型(VLM)提供了将动作与其视觉结果关联的共享推理空间。然而,根据完整动作序列判断完整长视频会产生冗长且嘈杂的上下文,其中短暂的局部动作证据可能被遗漏或稀释。我们提出WorldReward,这是一种基于VLM的成对偏好奖励模型,用于统一评估相机条件世界模型的动作一致性和视觉质量。WorldReward将成对视频分解为动作对齐的块,将每个块组织为结构化视觉证据,并通过投票将块级决策聚合为独立的视频级动作和视觉质量偏好。为训练该模型,我们构建了大规模推理增强偏好数据集,使用前沿VLM生成的结构化判断,并通过基于工具的智能体审计和针对性人工审核进行优化。我们还推出了WorldReward-Bench,这是一个人工标注的基准,用于衡量奖励模型在动作一致性、外观质量和动作质量三个维度上与人类偏好的一致性。WorldReward在所有三个维度上均达到最高一致性,分别超过GPT-5.5 3.42、1.45和3.56个百分点。当用于HY-WorldPlay 1.5的RL后训练时,它在短期到长期范围内持续提升动作执行和视觉质量。

英文摘要:

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.

补充信息

相关深度报道

↑