arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26821cs.RO

TemporalFlow-VLA:学习基于物理基础的执行历史以实现长 horizon 机器人操作

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

发表机构香港科技大学(广州) · 浙江大学 · 西蒙菲莎大学
另 1 家 · 查看机构详情
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Zhejiang University(浙江大学)
  • Simon Fraser University(西蒙菲莎大学)
  • AgiBot(云圣智能科技(上海)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Jiarui Yang, Yehao Lu, Yuning Su, Yu Zhong, Yufeng Xie, Yazhou Zhang, Haiyu Lan, Kaixiang Lu, Peiwen Lin, Chuang Wang, Junwei Liang, Enyu Li

首次发表
浏览论文内容

中文总结 AI 辅助

TemporalFlow-VLA 针对长 horizon 机器人操作中视觉相似状态需不同动作的问题,通过物理监督学习紧凑执行历史,在 LIBERO、RoboTwin 等任务上取得优异性能,且部署无额外开销。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型利用预训练的视觉-语言表示进行机器人控制,但仅添加历史帧无法可靠捕获近期的物理变化。在多阶段操作中,视觉相似的状态可能需要根据先前的执行采取不同的动作,这一问题尤为突出。为应对这一挑战,我们提出 TemporalFlow-VLA,该模型通过基于物理基础的时间监督学习紧凑的执行历史。利用记录的机器人状态、机器人几何结构和校准后的相机,我们构建机器人-表面时间流作为仅用于训练的目标,并监督两个与执行对齐的时间查询,这些查询为动作专家提供结构化历史。几何监督路径在部署时不进行评估。TemporalFlow-VLA 在 LIBERO 上实现了 97.63±0.26% 的平均成功率,其中 LIBERO Long 的成功率为 96.60±0.87%,在 12 项 RoboTwin 任务上的干净/随机成功率为 85.5%/84.2%。该模型在长 horizon、多阶段操作上相较于先前方法展现出最明显的优势。受控历史干预实验表明,动作预测同时依赖于历史内容和时间顺序。通过异步特征缓存,时间条件维持了单帧级别的服务器端采样延迟,且无额外的历史编码开销。总体而言,TemporalFlow-VLA 提供了一种紧凑、基于物理基础的接口,可用于利用有序的执行历史,且在部署时无需显式的运动估计或几何处理。

英文摘要

Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problematic in multi-stage manipulation, where visually similar states may require different actions depending on prior execution. To address this challenge, we present TemporalFlow-VLA, which learns compact execution history through physically grounded temporal supervision. Using recorded robot states, robot geometry, and calibrated cameras, we construct robot-surface temporal flow as a training-only target and supervise two execution-aligned temporal queries that provide structured history to the action expert. The geometric supervision path is not evaluated at deployment. TemporalFlow-VLA achieves 97.63 +/- 0.26% average success on LIBERO, including 96.60 +/- 0.87% on LIBERO Long, and 85.5%/84.2% Clean/Randomized success across 12 RoboTwin tasks. It shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation. Controlled history interventions show that action prediction depends on both historical content and temporal order. With asynchronous feature caching, temporal conditioning maintains single-frame-level server-side sampling latency without additional historical-encoding overhead. Overall, TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment.

↑