arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04964cs.AIcs.LG

WorldCycle:用于长时序视频世界模型的自验证强化学习

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对交互式视频世界模型的误差累积问题,提出自验证RL框架WorldCycle,利用可逆动作循环实现无标注监督,在CycleBench基准上大幅降低状态返回漂移并提升复合动作准确率。

中文摘要 AI 辅助

交互式视频世界模型是长时序规划与探索的核心,但存在误差累积问题。强化学习(RL)等训练后方法可改进这类模型,却遭遇验证瓶颈:对于任意动作序列,不存在真实未来状态来衡量长期漂移。核心洞见在于可逆动作循环可实现这种验证:由其逆序列构成的序列必然能解析地返回初始状态,从而为长时序正确性提供无标注监督。基于此,我们提出WorldCycle,一种自验证RL框架,它从普通动作序列构建闭合动作循环及其重复执行,并优化两种互补奖励:空间闭合奖励,强制正向与反向片段的对称性;时间一致性奖励,对齐循环重复执行时的状态。这些奖励迫使模型将动作学习为一致的状态算子,而非记忆的时间模式,且能自然扩展到基础模型处理不佳的分布外复合动作循环。我们还发布了CycleBench,用于评估复杂动作结构下状态返回能力的诊断基准。WorldCycle将状态返回漂移降低最多44%,复合动作准确率较基础模型提升近4倍,为物理驱动的世界模型提供了重要基础。

英文摘要

Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.

补充信息

↑