发表机构
New York University Abu Dhabi; Fudan University(纽约大学阿布扎比分校; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人形全身运动控制器在持续学习新技能时遗忘旧技能的问题,提出相似性引导的LoRA-PNN策略,通过两级运动相似性度量分配容量,实现强前向迁移和高效率,节省94.5%参数,并在Unitree G1上部署。
AI 中文摘要
人形全身控制器现在能够跟踪多种动态运动,但它们通常在离线状态下训练,然后被冻结,因此向这样的控制器教授新技能往往会侵蚀它已经掌握的技能。我们研究人形全身运动的持续学习,其中单个控制器必须从顺序任务流中获取技能,而不重新访问过去的数据。我们引入了相似性引导的LoRA-PNN,一种渐进式神经网络(PNN)策略,通过构造防止灾难性遗忘,同时通过轻量级低秩适应在技能间重用知识。一种基于动态时间规整并通过最优传输聚合的两级运动相似性度量,决定基于哪个先前技能以及分配多少新容量,从而产生强的前向迁移和大的效率提升。在六个顺序学习的技能类别中,我们的相似性引导的LoRA策略在所有方法中取得了最佳的前向迁移(0.125对0.079)和最高的平均准确率,同时节省了高达94.5%的可训练参数和40.8%的训练时间。由此产生的控制器达到了96.13%的模拟到模拟迁移,并部署在物理Unitree G1上。我们的代码可在https URL获取。
英文摘要
Humanoid whole-body controllers can now track a diverse set of dynamic motions, but they are typically trained offline and then frozen, so teaching such a controller a new skill tends to erode the skills it already mastered. We study continual learning for humanoid whole-body motion, where a single controller must acquire skills from a sequential task stream without revisiting past data. We introduce Similarity-guided LoRA-PNN, a progressive neural network (PNN) policy that prevents catastrophic forgetting by construction while reusing knowledge across skills through lightweight low-rank adaptation. A two-level motion-similarity measure, built from dynamic time warping aggregated by optimal transport, decides which prior skill to build on and how much new capacity to allocate, yielding strong forward transfer and large efficiency gains. Across six sequentially learned skill categories, our similarity-guided LoRA policy attains the best forward transfer (0.125 vs. 0.079) and the highest average accuracy among all methods, while saving up to 94.5% of trainable parameters and 40.8% of training time. The resulting controller reaches 96.13% sim-to-sim transfer and is deployed on a physical Unitree G1. Our code is available at https://anonymous.4open.science/r/continual-humanoid-learning-35D3.
Comments25 pages, 7 figures