发表机构
The University of Tokyo; ByteDance Seed; The University of Hong Kong; Shanghai Jiao Tong University; Tsinghua University(东京大学; 字节跳动 Seed; 香港大学; 上海交通大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SiMDex框架,通过三层流程从海量自我中心人类视频中挖掘与任务相关的子集,用于VLA后训练,仅用少量样本便显著提升了机器人灵巧操作的成功率,证明选择性数据整理更有效。
AI 中文摘要
近年来,以自我为中心的人类视频被大规模用于机器人操作的趋势呈爆炸式增长,但仍不清楚哪些数据真正有益于灵巧操作。我们提出了SiMDex,一种基于相似度的数据挖掘框架,将灵巧操作中用于VLA(视觉语言动作模型)后训练的人类数据选择问题转化为推荐问题。对于每个机器人演示,SiMDex采用三层召回-排序-重排序流程,从约3200万个自我中心人类样本池中提取与任务相关的子集,其在形态无关的动作空间中运行,无需修改VLA架构或训练。与使用等量随机采样人类数据训练的强基线相比,SiMDex仅使用约149万个挖掘样本(占池的不到5%),就将总体成功率从47.7%提升至61.1%,表明选择性数据整理优于无差别数据混合。
英文摘要
Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulation. We present SiMDex, a similarity-based data mining framework that casts human data selection for VLA post-training in dexterous manipulation as a recommendation problem. For each robot demonstration, SiMDex employs a three-layer recall-ranking-re-ranking pipeline to extract task-relevant subsets from a pool of ~32M egocentric human samples, operating in a morphology-agnostic action space that requires no changes to VLA architecture or training. Against a strong baseline trained with an equal amount of randomly sampled human data, SiMDex uses only ~1.49M mined samples (<5% of the pool) yet improves the overall success rate from 47.7% to 61.1%, showing that selective curation outperforms indiscriminate data mixing.
Comments12 pages, 4 figures. Project page: https://lin-nie.github.io/SiMDex/