arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoSteer:一个用于从第一人称视角视频实现可控灵巧操作的全栈系统

EgoSteer: An Open-Source Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

Yifan Zhong, Zhang Chen, Tianrui Guan, Fanlian Zeng, Ka Nam Lui, Yuyao Ye, Tingrui Zhang, Jiayi Li, Tianjia He, Wenjie Lou, Ruilin Yan, Xinhao Ji, Guangyu Zhao, Jiayuan Zhang, Wenxi Xu, Chengdong Ma, Yuanpei Chen, Yaodong Yang

arXiv 2607.09701首次发表:更新:

发表机构

Institute for AI, PKU; PKU-PsiBot Joint Lab; UPenn(北京大学人工智能研究院; 北大-灵犀机器人联合实验室; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对灵巧手系统因缺乏数据而无可控性的问题,提出全栈系统EgoSteer,集成EgoSmith数据管道等,通过人类数据预训练赋予语言引导操作先验,经机器人后训练强化,能执行多样任务,还开源相关内容。

AI 中文摘要

可控性是通用机器人策略的一项关键能力,但在灵巧手系统中却普遍缺失,因为缺乏大规模、语言对齐且动作准确的示范数据。为解决这一瓶颈,我们提出了一个全栈系统,该系统能从第一人称人类视频扩展灵巧VLA预训练,并实现数据高效的真实机器人后训练。它集成了EgoSmith数据管道,一个统一的机器人堆栈以及EgoSteer。人类数据预训练为EgoSteer配备了语言引导的操作先验知识,通过机器人后训练和DAgger优化得到强化。实验表明,EgoSteer能在40多个不同任务中稳健执行自由形式指令,预训练模型还能少样本适应复杂的长期任务。我们开源了系统、数据和模型。

英文摘要

The enduring vision of general-purpose robots serving humanity hinges fundamentally on policy steerability. However, prevailing paradigms of learning from expert demonstrations demand massive real-world data even on simplified grippers, rendering them prohibitively expensive for high-dimensional, data-scarce dexterous hands. To overcome this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9,606 hours of pre-training data with 8.3x higher throughput and better accuracy than prior SOTA; a unified Robot Stack for teleoperation and human-in-the-loop correction tailored for dexterous hands; and EgoSteer, a world-model-enhanced VLA operating on a morphology-aligned action space. Human data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and further refined via DAgger. Empirically, EgoSteer robustly executes free-form instructions across 45 diverse tasks, demonstrating adherence to user intent amid multiple candidate tasks and generalization. The pre-trained model also few-shot adapts to five complex long-horizon tasks, including box folding, on two embodiments with 79% average progress. All system code, datasets, model checkpoints, and an evaluation gallery are publicly available at https://egosteer.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑