发表机构
Shanghai Innovation Institute; Agibot(上海创新研究院; 智元机器人)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对灵巧操作中触觉信号稀疏的挑战,本文构建200小时双臂数据集并提出STAR训练方案,通过联合预训练与稀疏表示,在真实任务中达到61%平均成功率。
AI 中文摘要
灵巧操作需要协调的多指控制与有效的触觉反馈,然而,由于缺乏大规模真实世界数据以及难以从稀疏触觉信号中提取有效表示,学习这些能力仍然具有挑战性。我们构建了一个机器人平台和遥操作系统,收集了200小时的双臂灵巧操作数据集,其中包含同步的视觉、触觉和语言标注,涵盖65个任务中的10,576条轨迹,其中69.5%涉及灵巧的多指操作。我们进一步提出了STAR,一种针对视觉-触觉-语言-动作(VTLA)模型的集成训练方案,通过视觉-触觉联合预训练、稀疏全局触觉令牌表示和稀疏未来触觉预测来解决触觉信号在空间、时间和信息上的稀疏性。在该数据集上训练后,STAR在四个真实世界任务中实现了61%的平均成功率,每个任务使用100条训练后轨迹,展示了在任务特定训练后灵巧操作的性能。
英文摘要
Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.