发表机构
Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过分析不同课程调度的优化动力学,提出衡量跨难度知识迁移的Relative Transfer指标,据此推导TDCS方法,经多推理基准实验证明其性能优于代表性调度策略,为课程学习提供统一优化解释。
AI 中文摘要
课程学习已通过将训练数据按从易到难组织,被广泛应用于大语言模型的后训练中。然而,其在推理任务上的有效性差异极大,这表明不存在单一的通用最优课程,也引发了一个基本问题:什么决定了课程学习何时有效?本文通过分析不同课程调度策略诱导的优化动力学来回答该问题。研究表明,不同难度水平间的迁移关系刻画了课程学习诱导的优化动力学,进而解释了不同课程调度策略的有效性,并将该关系形式化为相对迁移(Relative Transfer),这是一种衡量跨难度知识迁移的原则性指标。基于该指标,本文推导了感知迁移的动态课程采样方法(Transfer-aware Dynamic Curriculum Sampling, TDCS),该方法在整个训练过程中根据估计的迁移关系动态调整采样分布。在多个推理基准上开展的大量实验表明,TDCS在不同任务、模型规模和训练范式下,始终优于代表性的调度策略。更重要的是,本文的工作通过跨难度迁移为课程学习提供了一种统一的、基于优化的解释。
英文摘要
Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? In this paper, we answer this question by analyzing the optimization dynamics induced by different curriculum schedules. We show that the transfer relationship between different difficulty levels characterizes the optimization dynamics induced by curriculum learning, which in turn explains the effectiveness of different curriculum schedules, and formalize this relationship as Relative Transfer, a principled measure of cross-difficulty knowledge transfer. Based on this measurement, we derive Transfer-aware Dynamic Curriculum Sampling (TDCS), which dynamically adjusts the sampling distribution according to the estimated transfer relationship throughout training. Extensive experiments on multiple reasoning benchmarks demonstrate that TDCS consistently outperforms representative scheduling strategies across different tasks, model scales, and training paradigms. More importantly, our work provides a unified optimization-based explanation of curriculum learning through cross-difficulty transfer.