WorldReward:面向相机条件世界模型的奖励建模
WorldReward: Reward Modeling for Camera-Conditioned World Models
- Fudan University(复旦大学)
- Tencent Hunyuan(腾讯混元)
- Shanghai Innovation Institute(上海创新研究院)
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出WorldReward,一种基于VLM的成对偏好奖励模型,解决相机条件世界模型现有奖励函数的缺陷,在WorldReward-Bench上超GPT-5.5,还提升了HY-WorldPlay 1.5的动作与视觉质量。
AI中文摘要:
相机条件世界模型生成交互式视频,要求指令动作能引发预期的场景变化,同时外观、几何结构和时间动态保持一致。现有奖励函数分别评估这些需求:基于几何的奖励估计轨迹执行情况,但无法判断执行动作的视觉质量;而基于图像的奖励测量帧质量,却无法捕捉动作执行或时间动态。我们假设视觉语言模型(VLM)提供了将动作与其视觉结果关联的共享推理空间。然而,根据完整动作序列判断完整长视频会产生冗长且嘈杂的上下文,其中短暂的局部动作证据可能被遗漏或稀释。我们提出WorldReward,这是一种基于VLM的成对偏好奖励模型,用于统一评估相机条件世界模型的动作一致性和视觉质量。WorldReward将成对视频分解为动作对齐的块,将每个块组织为结构化视觉证据,并通过投票将块级决策聚合为独立的视频级动作和视觉质量偏好。为训练该模型,我们构建了大规模推理增强偏好数据集,使用前沿VLM生成的结构化判断,并通过基于工具的智能体审计和针对性人工审核进行优化。我们还推出了WorldReward-Bench,这是一个人工标注的基准,用于衡量奖励模型在动作一致性、外观质量和动作质量三个维度上与人类偏好的一致性。WorldReward在所有三个维度上均达到最高一致性,分别超过GPT-5.5 3.42、1.45和3.56个百分点。当用于HY-WorldPlay 1.5的RL后训练时,它在短期到长期范围内持续提升动作执行和视觉质量。
英文摘要:
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.