arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构化参数环境下用于鲁棒导航策略的课程生成

Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

Prishita Ray

arXiv 2608.08545首次发表:更新:

发表机构

Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对结构化参数环境提出重参数化课程生成框架,结合分布偏移正则化,在OpenAI Gym的两类连续控制环境中,显著优于多种基线方法,提升了导航策略的鲁棒性。

AI 中文摘要

自主智能体的鲁棒导航策略必须在不断变化的环境条件(如转向速率、障碍物、摩擦力、坑洼和坡度)中实现泛化。课程生成通过逐步调整训练环境提供了一种提升泛化能力的原则性机制,但以样本高效且自动化的方式设计此类课程仍具挑战性。本文提出一种基于结构化连续环境参数的重参数化课程生成框架,采用单向基于梯度的优化方法。为提升由图像输入和标量输入组成的多模态观测空间中的鲁棒性,本文引入分布偏移正则化目标,以鼓励学习更细粒度的潜在表示。所提方法在两个连续控制类OpenAI Gym环境中进行评估:基于2D障碍物的赛车变体和双足机器人行走变体,其中耦合的环境参数共同影响策略性能。在5个随机种子下,本文方法始终优于普通策略训练、随机参数采样、人工课程、基于前沿的方法、自步强化学习(SPRL)、带高斯混合模型的绝对学习进度(ALP-GMM)以及反向课程学习基线。 ablation研究进一步证明了重参数化课程机制在两种环境中的有效性,同时凸显了辅助正则化目标对不同环境的差异化增益。

英文摘要

Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, friction, pits, and slopes. Curriculum generation provides a principled mechanism for improving generalization by progressively adapting training environments, but designing such curricula in a sample-efficient and automated manner remains challenging. This paper proposes a reparameterized curriculum generation framework for structured continuous environment parameters using unidirectional gradient-based optimization. To improve robustness in multimodal observation spaces consisting of image-based and scalar inputs, a distribution-shift regularization objective is incorporated to encourage the learning of finer-grained latent representations. The proposed method is evaluated across two continuous-control OpenAI Gym environments: a 2D obstacle-based Car Racing variant and Bipedal Walker variant, where coupled environment parameters jointly influence policy performance. Across five random seeds, our method consistently outperforms vanilla policy training, random parameter sampling, manual curricula, frontier-based methods, Self-Paced Reinforcement Learning (SPRL), Absolute Learning Progress with Gaussian Mixture Models (ALP-GMM), and reverse curriculum learning baselines. Ablation studies further demonstrate the effectiveness of the reparameterized curriculum mechanism across both environments, while highlighting environment-dependent benefits of the auxiliary regularization objective.

Comments22 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑