发表机构
South China University of Technology; The Hong Kong University of Science and Technology; Northwestern Polytechnical University; The Hong Kong Polytechnic University(华南理工大学; 香港科技大学; 西北工业大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人操作中历史依赖决策问题,提出多源数据集GiT,涵盖18个双手任务及反事实任务对,评估VLA模型,发现显著改进空间。
AI 中文摘要
机器人操作通常需要在当前观察不足以确定适当动作时,从过去的交互中推断与任务相关的状态。尽管在记忆增强的视觉-语言-动作(VLA)模型的基准测试方面取得了进展,但需要历史依赖语义推理的应用导向任务仍然代表性不足。我们引入了GiT(扎根于时间),一个用于在生物实验室、家庭和工业场景中基于过去事件进行操作决策定位的数据集和基准。它包含真实机器人和通用操作接口(UMI)风格的演示,涵盖18个双手任务,以及模拟数据和基于ManiSkill的评估套件,涵盖9个任务。细粒度的子任务注释和注释的反事实任务对(其中相似当前观察根据先前事件需要不同动作)支持策略学习和历史使用的针对性评估。在模拟和选定的真实世界任务上对代表性端到端VLA模型的评估揭示了历史依赖操作方面的显著改进空间。数据集和基准可在项目页面获取。
英文摘要
Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent semantic inference remain underrepresented. We introduce GiT (Grounded in Time), a dataset and benchmark for grounding manipulation decisions in past events across biolaboratory, household, and industrial scenarios. It includes real-robot and Universal Manipulation Interface (UMI) style demonstrations covering 18 bimanual tasks, together with simulation data and a ManiSkill-based evaluation suite covering nine tasks. Fine-grained subtask annotations and annotated counterfactual task pairs, in which similar current observations require different actions depending on prior events, support policy learning and targeted evaluation of history use. Evaluations of representative end-to-end VLA models in simulation and on selected real-world tasks reveal substantial room for improvement in history-dependent manipulation. The dataset and benchmark are available at the project page.