TemporalFlow-VLA:学习基于物理基础的执行历史以实现长 horizon 机器人操作
TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation
另 1 家 · 查看机构详情
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Zhejiang University(浙江大学)
- Simon Fraser University(西蒙菲莎大学)
- AgiBot(云圣智能科技(上海)有限公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
TemporalFlow-VLA 针对长 horizon 机器人操作中视觉相似状态需不同动作的问题,通过物理监督学习紧凑执行历史,在 LIBERO、RoboTwin 等任务上取得优异性能,且部署无额外开销。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型利用预训练的视觉-语言表示进行机器人控制,但仅添加历史帧无法可靠捕获近期的物理变化。在多阶段操作中,视觉相似的状态可能需要根据先前的执行采取不同的动作,这一问题尤为突出。为应对这一挑战,我们提出 TemporalFlow-VLA,该模型通过基于物理基础的时间监督学习紧凑的执行历史。利用记录的机器人状态、机器人几何结构和校准后的相机,我们构建机器人-表面时间流作为仅用于训练的目标,并监督两个与执行对齐的时间查询,这些查询为动作专家提供结构化历史。几何监督路径在部署时不进行评估。TemporalFlow-VLA 在 LIBERO 上实现了 97.63±0.26% 的平均成功率,其中 LIBERO Long 的成功率为 96.60±0.87%,在 12 项 RoboTwin 任务上的干净/随机成功率为 85.5%/84.2%。该模型在长 horizon、多阶段操作上相较于先前方法展现出最明显的优势。受控历史干预实验表明,动作预测同时依赖于历史内容和时间顺序。通过异步特征缓存,时间条件维持了单帧级别的服务器端采样延迟,且无额外的历史编码开销。总体而言,TemporalFlow-VLA 提供了一种紧凑、基于物理基础的接口,可用于利用有序的执行历史,且在部署时无需显式的运动估计或几何处理。
英文摘要
Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problematic in multi-stage manipulation, where visually similar states may require different actions depending on prior execution. To address this challenge, we present TemporalFlow-VLA, which learns compact execution history through physically grounded temporal supervision. Using recorded robot states, robot geometry, and calibrated cameras, we construct robot-surface temporal flow as a training-only target and supervise two execution-aligned temporal queries that provide structured history to the action expert. The geometric supervision path is not evaluated at deployment. TemporalFlow-VLA achieves 97.63 +/- 0.26% average success on LIBERO, including 96.60 +/- 0.87% on LIBERO Long, and 85.5%/84.2% Clean/Randomized success across 12 RoboTwin tasks. It shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation. Controlled history interventions show that action prediction depends on both historical content and temporal order. With asynchronous feature caching, temporal conditioning maintains single-frame-level server-side sampling latency without additional historical-encoding overhead. Overall, TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment.