arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越预测:利用回顾性世界建模引导VLM智能体

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo

arXiv 2609.39101首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; The Hong Kong Polytechnic University; China University of Geoscience(香港科技大学; 香港理工大学; 中国地质大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出回顾性世界建模范式,通过估计状态转换的回顾性归因分布并构建自一致性奖励,引导VLM智能体产生物理一致且任务有效的行动,显著提升策略鲁棒性与泛化能力。

AI 中文摘要

为VLM智能体配备世界建模能力在复杂推理和长期规划方面展现出强大潜力,同时减少策略学习对昂贵真实世界交互的依赖。现有方法主要依赖前瞻性模拟来预测候选行动的后果。然而,这种仅向前看的范式关注接下来会发生什么,对验证行动是否与观察到的状态转换因果一致提供的约束有限,可能导致看似合理但物理上不连贯的行为。在本文中,我们挑战将世界建模仅视为前瞻性预测的观点,引入回顾性世界建模(Retrospective World Modeling),一种新的智能体学习范式,使智能体能够通过估计给定转换最可能引发行动的回顾性归因分布$P(\hat{a}_{t}|s_t, s_{t+1})$来进行向后推理。基于此能力,我们提出自一致性奖励(Self-Consistency Reward, SCR),一种衡量策略行动与回顾性解释之间概率一致性的内在信号。将SCR整合到强化学习中,提供密集的转换级反馈,引导智能体朝向既任务有效又物理基础扎实的行为。跨多种智能体任务的广泛实验表明,我们的方法在策略鲁棒性和泛化性上显著优于仅前瞻性世界建模基线。

英文摘要

Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on what will happen next and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution $P(\hat{a}{t}|s_t, s{t+1})$ for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.

CommentsAccepted at NeurIPS 2026 (27 pages, 8 figures)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑