SMART:通过大规模合成预训练实现关节物体的零样本仿真到现实操作
SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
浏览论文内容
中文总结 AI 辅助
针对关节物体操作中真实演示稀缺的问题,提出SMART系统,通过仿真平台和智能体任务生成合成超百万演示,预训练VLA模型实现零样本仿真到现实迁移。
中文摘要 AI 辅助
与关节物体交互的能力对于具身智能系统至关重要,但由于涉及精确的接触和遵循约束的运动,收集大规模真实世界演示仍然具有挑战性。尽管仿真提供了一个有前景的替代方案,现有的合成数据工作覆盖的关节物体类别有限,而通用合成流水线缺乏对部件级语义和关节约束的显式设计,阻碍了智能体任务生成和高质量关节操作演示的可扩展合成。为弥合这一差距,我们引入SMART,一个利用大规模合成操作演示进行关节物体操作的可扩展系统。其核心是,我们开发了SMART-Sim,一个具有关节感知设计的仿真平台,能够实现有效的任务生成和高效的演示收集。基于SMART-Sim,我们应用智能体任务生成并设计了一个可扩展的分布式合成系统,利用它们合成SMART-Data,包含超过100万条演示,涵盖44种原子任务类型、5种机器人设置和2,507个关节物体。在SMART-Data上预训练的视频-语言-动作(VLA)模型在仿真基准上表现出有竞争力的性能,并在真实世界关节物体操作任务中实现了零样本仿真到现实迁移和可扩展性能。这突显了合成演示在提供有效和可扩展监督以提升VLA模型在接触丰富的关节物体操作中性能的潜力。
英文摘要
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
发表机构
- Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)
- Technical University of Munich(慕尼黑工业大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Shanghai Jiao Tong University(上海交通大学)
- Tsinghua University(清华大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。