发表机构
National Key Laboratory of General Artificial Intelligence, BIGAI; Beijing Normal University; Communication University of China; Shanghai Jiao Tong University(通用人工智能国家重点实验室(BIGAI); 北京师范大学; 中国传媒大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对非结构化地形四足机器人运动,提出轨迹级自动课程学习框架,通过生成地形轨迹任务形成闭环课程,显著提升了运动策略的轨迹成功率。
AI 中文摘要
针对复杂非结构化地形训练运动策略需要课程来避免早期探索失败。然而,由于非结构化地形缺乏课程设计所需的明确难度排序,现有方法采用基于参数化地形的启发式课程。这种抽象限制了泛化能力,因为策略可能过度适应接近固定的感知模式。为解决该问题,我们提出了OURNAME,一种轨迹级自动课程学习(Trajectory-level Automatic Curriculum Learning)框架,可直接从非结构化地形地图生成训练任务。在每次课程更新时,评估器学习当前策略的难度函数,将给定轨迹任务映射为难度分数;采样器随后在学习到的评估器引导下提出新轨迹,作为下一次策略更新的课程。这形成了闭环,使课程与不断演进的策略迭代匹配。定量与定性实验显示,OURNAME在非结构化地形上持续提供有效课程,相比无课程的直接训练,轨迹成功率提升56.3%;与手工设计的课程学习相比,在最难地形任务上成功率提升18.5%,在相同障碍物类型下从不同接近方向评估 traversal 时提升最高达39.74%。
英文摘要
Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures. However, since unstructured terrain lacks explicit difficulty ordering for curriculum design, existing methods resort to heuristic curricula over parameterized terrains. This abstraction limits generalization, as policies can overadapt to near-fixed perceptual patterns. To address this, we propose \textbf{\ourname{}}, an \textbf{T}rajectory-level \textbf{A}utomatic \textbf{C}urriculum \textbf{L}earning framework that generates training tasks directly from unstructured terrain maps. At each curriculum update, the evaluator learns a difficulty function for the current policy that maps a given trajectory task to a difficulty score. The sampler then proposes new trajectories guided by the learned evaluator as the curriculum for the next policy update. This forms a closed loop in which the curriculum is iteratively matched to the evolving policy. Quantitative and qualitative experiments show that \ourname{} continuously provides effective curricula on unstructured terrain, improving trajectory success rate by \(56.3\%\) over direct training without curriculum. Compared with handcrafted curriculum learning, our method improves success rate by \(18.5\%\) on the hardest terrain tasks and by up to \(39.74\%\) when evaluating traversal from diverse approach directions on the same obstacle type.