arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16864cs.ROcs.CVcs.LG

TEMPO:学习动态机器人操作的时间上下文

TEMPO: Learning Temporal Context for Dynamic Robot Manipulation

Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型在动态操作中因缺乏时间上下文而失败的问题,提出TEMPO,通过运动摘要和本体感觉历史两种时间输入增强预训练模型,显著提升动态任务成功率并解决状态混叠。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在准静态操作中取得了显著性能,但在动态操作任务中表现不佳,因为它们在推理时仅基于单一观测进行操作。我们识别出导致这一局限性的两种表征失败。第一种是运动模糊性,即单一观测不包含场景动态,因此无法预测移动物体的未来状态。第二种是状态混叠,即任务中不同时间点的视觉上相似的观测需要不同的动作。我们认为这些失败无论模型规模或推理延迟如何都会持续存在,表明瓶颈在于缺失时间上下文而非模型容量。基于这一洞察,我们提出TEMPO,它通过两种时间输入增强预训练的VLA:从冻结的视频基础模型中提取的运动摘要以解决运动模糊性,以及紧凑的本体感觉历史以解决状态混叠。TEMPO无需修改主干网络,且在训练或部署时仅增加极小的计算开销。在四个动态操作任务中,它将瓶子交接成功率从44%提升至74%,并且是唯一解决状态混叠的方法。探针和消融研究证实,每个时间信号独立地解决其对应的失败。我们进一步发布TEMPO-Bench,一个包含超过5万标注帧的基准,用于以回归和多项选择格式评估运动感知的机器人感知。项目网站:此https URL

英文摘要

Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at inference time. We identify two representational failures that underlie this limitation. The first is motion ambiguity, where a single observation does not include scene dynamics and therefore cannot anticipate the future state of moving objects. The second is state aliasing, where visually similar observations from different points in a task require different actions. We argue that these failures persist regardless of model scale and inference latency, showing that the bottleneck is missing temporal context rather than model capacity. Based on this insight, we propose TEMPO, which augments a pretrained VLA with two temporal inputs: a motion summary extracted from a frozen video foundation model to resolve motion ambiguity and a compact proprioceptive history to resolve state aliasing. TEMPO requires no modification to the backbone and adds minimal compute overhead at training or deployment. Across four dynamic manipulation tasks, it improves Bottle Handover success from 44% to 74% and is the only method that solves state aliasing. Probing and ablation studies confirm that each temporal signal independently addresses its corresponding failure. We further release TEMPO-Bench, a benchmark of over 50k annotated frames for evaluating motion-aware robot perception in both regression and multiple-choice formats. Project Website: https://tempo-robot.github.io/

发表机构

  • University of California, Irvine(加利福尼亚大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑