InstantMimic:一个在数秒内学习基于物理的技能的高性能系统
InstantMimic: A High Performance System for Learning Physics-based Skills in Seconds
AI总结:
InstantMimic通过GPU原生统一流水线集成模拟、环境计算、策略推理与更新,将基于物理的技能训练时间缩短至数秒,并支持LLM智能体驱动的超参数搜索。
AI中文摘要:
基于物理的角色控制是计算机图形学和机器人学中长期存在的挑战,需要满足复杂动力学并产生逼真运动的策略。最近的深度强化学习方法,特别是诸如DeepMimic等模仿学习方法,其影响已超越动画领域,通过实现敏捷且富有表现力的行为而影响机器人学。尽管这些方法取得了令人印象深刻的结果,但在实践中训练时计算效率仍然低下。尽管使用了GPU加速模拟,我们发现端到端流水线常常因物理求解器之外的额外开销而未能充分利用硬件,这些开销由碎片化的GPU内核和关键路径中的CPU内存访问引起。我们提出了InstantMimic,一个通过使整个训练循环GPU原生化来解决这些低效问题的系统。基于GPU原生的物理后端,我们的统一流水线在单一执行流中集成了模拟、环境计算、策略推理和策略更新。因此,InstantMimic将多种基于物理的技能的训练时间缩短至数秒,并使基于LLM智能体的超参数搜索变得实用。
英文摘要:
Physics-based character control is a long-standing challenge in computer graphics and robotics, requiring policies that satisfy complex dynamics while producing realistic motion. Recent Deep RL approaches, particularly imitation learning methods such as DeepMimic, have had broad impact beyond animation, influencing robotics by enabling agile and expressive behaviors. While these approaches achieve impressive results, they remain computationally inefficient to train in practice. Despite GPU-accelerated simulation, we find that end-to-end pipelines often underutilize hardware due to overheads outside the physics solver, caused by fragmented GPU kernels and CPU memory access in the critical path. We present InstantMimic, a system that addresses these inefficiencies by making the entire training loop GPU-native. Built on a GPU-native physics backend, our unified pipeline integrates simulation, environment computation, policy inference, and policy updates within a single execution flow. As a result, InstantMimic reduces training time for diverse physics-based skills to a few seconds and makes LLM-agent-driven hyperparameter search practical.