发表机构
Deemos Corporation; ShanghaiTech University; University of Cambridge; D-Robotics(Deemos 公司; 上海科技大学; 剑桥大学; D-Robotics 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LACE-CRAFT通过策略继承和黑板协作,在机器人协同设计中实现形态搜索与策略学习的结合,显著提升多基准性能,并支持从仿真到硬件的流程。
AI 中文摘要
机器人协同设计将形态搜索与策略学习相结合,然而从头训练每一个新设计会丢弃已获得的控制经验。我们提出了LACE-CRAFT,它比较了在当前机器人上的持续学习与对新形态-奖励对的策略适应。LACE恢复现有智能体的完整学习状态,并用其智能体参数和观测统计初始化兼容的挑战者。一个固定的任务指标在两个分支和冻结的现有智能体之间进行选择。CRAFT通过共享的实验记录和行为重放来协调反馈、形态、奖励和集成角色,以生成并交叉审查配对提案。一个生成式扩展将生成的网格转换为可编辑的铰接模型,并配置关节、执行器接口和一致更新的仿真资产。在五个运动基准上,三个评估种子的平均得分比D2C高6.4%至91.9%。两种方法在五轮中训练30个新的形态-奖励对;LACE额外使用四个持续训练单元。五项任务的消融研究考察了策略继承和重放衍生的反馈。一个制造的样机演示了室内行走,并展示了从几何到硬件的流程。
英文摘要
Robot co-design couples morphology search with policy learning, yet training every new design from scratch discards acquired control experience. We present LACE-CRAFT, which compares continued learning on the current robot with policy adaptation to new morphology-reward pairs. LACE resumes the incumbent's full learning state and initializes compatible challengers with its actor parameters and observation statistics. A fixed task metric selects among both branches and the frozen incumbent. CRAFT coordinates Feedback, Morphology, Reward, and Integration roles through shared experimental records and behavioral replays to generate and cross-review paired proposals. A generative extension converts generated meshes into editable articulated models with configured joints, actuator interfaces, and consistently updated simulation assets. Across five locomotion benchmarks, mean scores over three evaluation seeds are 6.4-91.9% higher than D2C. Both methods train 30 new morphology-reward pairs over five rounds; LACE additionally uses four continuation training units. Five-task ablations examine policy inheritance and replay-derived feedback. A fabricated prototype demonstrates indoor walking and illustrates the geometry-to-hardware workflow.
Comments16 pages including appendix. Project website: https://deemostech.github.io/lace-craft/