发表机构
University of Cambridge; Industrial Next(剑桥大学; Industrial Next)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过真实到仿真再到真实的协同训练,区分并独立调控世界接地与行为接地,在灵巧抓取分拣任务中将成功率从52%提升至86%,揭示了两者的互补作用及对基础模型协同训练的益处。
AI 中文摘要
仿真可以扩充稀缺的真实演示数据以用于协同训练,然而世界保真度以及与人类行为的相似性如何影响策略性能仍不清楚。我们区分了世界接地(将仿真与真实系统对齐)和行为接地(将仿真轨迹与人类运动对齐)。我们构建了一个真实到仿真再到真实的流水线,独立地变化这些轴以生成用于协同训练的数据。在一个动态灵巧抓取与分拣任务中,完全接地的协同训练将成功率从52%提升至86%;跨配置平均来看,世界接地将成功率提高了18个百分点,行为接地提高了10个百分点。部署的策略表现得像真实衍生与仿真衍生策略的混合体,在覆盖状态下模仿真实演示,在其他情况下依赖仿真行为,我们通过潜在空间分析对此进行了考察。这些结果共同表明互补作用:世界接地使策略能够利用超出真实数据覆盖范围的仿真经验,而行为接地主要在世界接地不完善时发挥作用。在协同训练基础模型时,接地仿真仍然有益。
英文摘要
Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. We distinguish world grounding, which aligns simulation with the real system, and behavior grounding, which aligns simulated trajectories with human motion. We build a real2sim2real pipeline that varies these axes independently to generate data for co-training. On a dynamic dexterous pick-and-sort task, fully grounded co-training raises success from 52% to 86%; averaged across configurations, world grounding improves success by 18 percentage points and behavior grounding by 10. Deployed policies behave like a mixture of real-derived and simulation-derived policies, imitating real demonstrations in covered states and relying on simulated behavior elsewhere, which we examine through latent-space analysis. Together, these results suggest complementary roles: world grounding lets policies use simulated experience beyond real-data coverage, while behavior grounding matters mainly when world grounding is imperfect. Grounded simulation remains beneficial when co-training foundation models.
CommentsProject Website: https://industrialnext.github.io/r2s2r-grounding/