发表机构
The Hong Kong University of Science and Technology (Guangzhou); Xiaomi Robotics; Southeastern University; Tsinghua University; Zhejiang University; Westlake University; Symbiosis Robotics; Peking University; Xi’an Jiaotong University; Beijing Academy of Artificial Intelligence(香港科技大学(广州); 小米机器人; 东南大学; 清华大学; 浙江大学; 西湖大学; 共生机器人; 北京大学; 西安交通大学; 北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EmbodiedSmith通过递归自我改进循环统一场景与任务生成,扩展具身数据,支持多种机器人形态,实验证明其提升数据质量、多样性与泛化能力。
AI 中文摘要
扩展机器人基础模型需要多样化的训练数据和可靠的评估环境。仿真提供了一种可扩展的解决方案,但现有的生成流程仍受限于预定义的资产和技能、场景生成与任务生成之间的脱节,以及对复杂具身形态和物理的有限支持。我们提出了EmbodiedSmith,一个通过递归自我改进(RSI)实现可扩展具身数据生成的框架。EmbodiedSmith在一条流程中统一了资产、场景和任务生成,支持自主创建和语言驱动的定制。其核心是一个智能体驱动的优化循环:场景生成预测下游任务需求,而任务生成引导针对性的场景编辑,使场景和任务能够迭代地相互改进。这种联合优化提高了任务生成的成功率,包括长时程任务。该框架还支持移动操作器、人形机器人和灵巧手,以及涉及可变形物体和流体的交互,拓宽了生成数据中所代表的行为和物理现象的范围。这些能力共同提供了一个灵活的仿真引擎,用于机器人预训练和评估。大量实验验证了所生成数据的质量、多样性和生成效率,而下游策略实验表明,数据多样性的增加提高了泛化能力。
英文摘要
Scaling robotic foundation models requires diverse training data and reliable evaluation environments. Simulation offers a scalable solution, yet existing generation pipelines remain constrained by predefined assets and skills, a disconnect between scene generation and task generation, and limited support for complex embodiments and physics. We introduce EmbodiedSmith, a framework for scalable embodied data generation through recursive self-improvement (RSI). EmbodiedSmith unifies asset, scene, and task generation in a pipeline that supports autonomous creation and language-driven customization. Its core is an agentic refinement loop: scene generation anticipates downstream task requirements, while task generation guides targeted scene edits, allowing scenes and tasks to iteratively improve one another. This joint refinement improves task generation success, including for long-horizon tasks. The framework further supports mobile manipulators, humanoids, and dexterous hands, as well as interactions involving deformable objects and fluids, broadening the range of behaviors and physical phenomena represented in generated data. Together, these capabilities provide a flexible simulation engine for both robot pretraining and evaluation. Extensive experiments validate the quality, diversity, and generation efficiency of the resulting data, while downstream policy experiments demonstrate that increased data diversity improves generalization.