arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CST-WM:用于具身视觉跟踪的因果结构世界模型

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang

arXiv 2609.06302首次发表:更新:

发表机构

New York University Abu Dhabi(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对具身视觉跟踪中的因果幻觉问题,提出因果结构世界模型CST-WM,通过阻断动作到目标证据的直接路径,结合模型预测控制,在EVT-Bench和Habitat 3.0上提升跟踪与重获取性能。

AI 中文摘要

具身视觉跟踪要求机器人不仅对当前视图做出反应,还要选择能够在自我运动、遮挡和干扰物下保持或恢复移动目标未来证据的动作。因此,这是一个关于未来目标可观测性和表观尺度的预测性决策问题。一个核心困难是任务特定的因果幻觉形式:在动作条件预测中,模型可以利用机器人控制与目标相关观测之间的强相关性,通过幻觉当前动作到目标证据的直接因果效应,而不是让动作仅通过机器人运动和由此产生的观测变化来影响该证据。这种捷径产生了具有错误语义的看似合理的未来,不利于面向跟踪的规划和重新获取。我们提出了CST-WM,一种因果结构世界模型,它将潜在状态分解为目标证据、机器人和观测分支,并对转移进行因式分解,使得直接动作注入到目标证据分支被阻断,而动作仍可用于机器人运动和观测更新。结合基于展开的模型预测控制,CST-WM在一个规划框架内支持稳定跟踪和临时目标重新获取。在EVT-Bench和Habitat 3.0上,涵盖标准跟踪、目标丢失恢复和跨数据集迁移,它相对于反应式基线和世界模型基线提高了跟踪质量、距离范围控制、安全性和重新获取能力;离线诊断显示更好的多步展开保真度、更强的规划价值一致性,以及显著减少的直接动作泄漏。对于具身视觉跟踪,仅靠未来预测是不够的:预测结构本身必须与目标证据如何进入规划相一致。

英文摘要

Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data, however, the behavior policy's actions are correlated with where the target is, so a generic predictor can learn a shortcut: it writes the current action directly into its prediction of target evidence, instead of letting the action affect that evidence only by moving the robot and changing what it observes. We call this failure causal hallucination; the resulting rollouts look plausible but rank candidate actions for the wrong reason. We propose CST-WM, a causally structured world model whose state is split into target-evidence, robot, and observation branches. Its transition removes the same-step edge from action to target evidence but keeps the path through robot motion and the resulting views, so candidate actions are still distinguished by their predicted ego-motion. With rollout-based model-predictive control, a single model handles both steady following and re-acquisition after target loss. On EVT-Bench and Habitat 3.0, covering standard tracking, target-loss recovery, and cross-dataset transfer, CST-WM improves following, distance-range control, safety, and re-acquisition over reactive trackers and world-model baselines, and removing the action mask causes the largest drop in re-acquisition among our ablations. Offline, CST-WM has lower multi-step rollout error, and its ranking of candidate actions agrees better with the simulator's. On a Unitree Go2 quadruped, CST-WM succeeds in 20 of 30 real-world trials under occlusion, distractor crossing, and fast motion, against 14 for TrackVLA.

Comments21 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑