发表机构
Princeton University; Microsoft Research; Harvard University(普林斯顿大学; 微软研究院; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于Wasserstein测地线的传输框架解耦课程学习因素,在合成套件上发现课程效应依赖上下文,从易到难排序提升困难层级表现。
AI 中文摘要
课程学习受若干耦合设计选择的支配——难度如何定义、样本如何排序、每个层级获得多少曝光量、以及训练在层级间移动的速度——这使得很难分离出真正起作用的因素。我们提出了Wasserstein课程路径,一个基于传输的简单框架,通过将课程表示为离散难度层级上训练分布的轨迹来解耦这些因素。在一个包含12个任务和33个难度轴的校准合成套件中,我们使用该框架在固定训练预算下分离排序、匹配曝光、端点平滑性和节奏的影响。我们发现课程效应强烈依赖于上下文:没有单一策略在任务、难度轴和预算上占主导地位,课程主要改变固定预算最有效花费的位置。在该框架内,从易到难的排序相对于曝光匹配的静态采样提高了困难层级的表现,表明其益处不能仅由累积曝光解释。我们进一步表明,端点平滑性和节奏显著影响课程在难度谱上有效的范围。最后,我们展示相同的传输视角自然支持通过几何学习节奏的扩展,以及超越一维排序的结构化难度空间的扩展。
英文摘要
Curriculum learning is governed by several coupled design choices---how difficulty is defined, how examples are ordered, how much exposure each level receives, and how quickly training moves across levels---making it hard to isolate what actually helps. We present Wasserstein curriculum paths, a simple transport-based framework that decouples these factors by representing curricula as trajectories of training distributions over discrete difficulty levels. Across a calibrated synthetic suite with 12 tasks and 33 difficulty axes, we use this framework to isolate the effects of ordering, matched exposure, endpoint smoothness, and pacing under fixed training budgets. We find that curriculum effects are strongly context-dependent: no single strategy dominates across tasks, difficulty axes, and budgets, and curricula mainly change where a fixed budget is spent most effectively. Within this framework, easy-to-hard ordering improves hard-level performance relative to exposure-matched static sampling, showing that the benefit is not explained by cumulative exposure alone. We further show that endpoint smoothness and pacing substantially affect where along the difficulty spectrum a curriculum is effective. Finally, we show that the same transport view naturally supports extensions to learned pacing through geometry and to structured difficulty spaces beyond one-dimensional orderings.
CommentsCOLM 2026