发表机构
The University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对强化学习中环境的有向无环课程图结构,提出主动课程学习框架PATH,通过采样多样化课程路径并重新分配训练资源,提升了模型的鲁棒性与泛化能力。
AI 中文摘要
在许多强化学习(RL)领域中,环境之间存在前提关系,例如难度递增的编辑或参数增量,这会诱导出有向无环课程图(DAG)。尽管这种结构通常仅被隐式利用,但显式建模它可以改善训练效果。我们提出了PATH,这是一种对课程图进行主动学习的课程学习框架。PATH首先通过采样多样化的课程路径来扩大覆盖范围,然后将训练资源重新分配到仍未掌握的区域。在多种环境中进行的实验表明,PATH通过显式利用图结构实现了出色的鲁棒性和泛化能力。
英文摘要
In many reinforcement learning (RL) domains, environments are connected by prerequisite relations, such as difficulty-increasing edits or parameter increments, which induce a directed acyclic curriculum graph (DAG). Although this structure is often exploited only implicitly, explicitly modeling it can improve training. We introduce PATH, a curriculum-learning framework that performs active learning over the curriculum graph. PATH first expands coverage by sampling diverse curriculum paths and then reallocates training toward regions that remain unmastered. Experiments across diverse environments show that PATH explicitly leverages the graph structure to achieve strong robustness and generalization.
Journal refProceedings of the International Conference on Machine Learning 2026