arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24101cs.RO

TrAct:通过视觉轨迹连接机器人控制与视觉预测

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

Zhi Cao, Howard Ji, Kevin Zhang, Kuangzhi Ge, Li Fei-Fei, Jiajun Wu, Huang Huang

首次发表
浏览论文内容

中文总结 AI 辅助

TrAct是基于世界模型的机器人决策框架,以视觉轨迹为控制与预测的中间接口,在LIBERO-INTEGRAL基准和Franka任务上,相较基线π₀.5显著提升了操作成功率与视频预测质量。

中文摘要 AI 辅助

机器人动作本质上是与具体实体绑定的,且与图像空间的视觉变化仅存在弱关联,这限制了其作为机器人世界模型条件信号的有效性。相比之下,视觉轨迹提供了一种与实体无关的表示,能体现任务相关点在场景中的移动方式,为准确且空间精确的未来视频预测提供密集的图像空间指导。基于这一观察,我们提出TrAct,一种基于世界模型的机器人决策框架,它将视觉轨迹用作控制与预测之间的中间接口。TrAct包含三个组件:一是视觉-语言-动作-轨迹模型(VLAT),它能从当前观测和语言指令中联合预测候选动作及对应的视觉轨迹;二是轨迹条件世界模型(TWM),它能在给定所提出轨迹的条件下预测未来视觉结果;三是视觉-语言奖励模型(VLAC),它能对预测结果进行评分。在推理阶段,VLAT生成候选动作-轨迹对,TWM推演其视觉后果,VLAC则选择预测结果最符合指令的轨迹,机器人随后执行与所选轨迹配对的动作。在提出的LIBERO-INTEGRAL基准和真实世界Franka操作任务上的实验显示,与强大的VLA基线π₀.₅相比,TrAct在模拟任务中将成功率从27%提升至55%,在真实世界任务中从49%提升至76%。此外,TWM相比动作条件世界模型(AWM)持续提升了视频预测质量。这些结果表明,视觉轨迹为机器人控制与视觉预测提供了有效的共享接口,可实现更准确的世界建模和更强的机器人泛化能力。

英文摘要

Robot actions are inherently embodiment-specific and only weakly aligned with image-space visual changes, limiting their effectiveness as conditioning signals for robot world models. In contrast, visual tracks provide an embodiment-agnostic representation of how task-relevant points move through a scene, offering dense image-space guidance for accurate and spatially precise future video prediction. Building on this observation, we propose TrAct, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction. TrAct consists of three components: a Vision-Language-Action-and-Track model (VLAT) that jointly predicts candidate actions and corresponding visual tracks from the current observation and language instruction; a track-conditioned world model (TWM) that predicts future visual outcomes conditioned on the proposed tracks; and a vision-language reward model (VLAC) that scores the predicted outcomes. At inference time, VLAT generates candidate action-track pairs, TWM rolls out their visual consequences, and VLAC selects the track whose predicted outcome best satisfies the instruction; the action paired with the selected track is then executed by the robot. Experiments on the proposed LIBERO-INTEGRAL benchmark and real-world Franka manipulation show that TrAct improves success rates from 27% to 55% in simulation and from 49% to 76% on real-world tasks compared with the strong VLA baseline $π_{0.5}$. Furthermore, TWM consistently improves video prediction quality over the action-conditioned world model (AWM). These results demonstrate that visual tracks provide an effective shared interface between robot control and visual prediction, enabling more accurate world modeling and stronger robot generalization.

发表机构

  • University of Michigan(密歇根大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑