arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35426cs.LGcs.AIcs.CL

前沿学习:在能力边缘训练LLM推理器

Frontier Learning: Training LLM Reasoners at the Edge of Capability

  • University of Basel(巴塞尔大学)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi

AI总结:

提出前沿学习,一种开放式后训练方法,通过程序化生成器在线产生前沿难度问题,利用遗憾信号探索能力边缘,持续提升LLM推理性能。

AI中文摘要:

基于强化学习的大语言模型(LLM)后训练已成功应用于提升其推理能力。现有流程主要使用GRPO损失,在训练前指定的固定问题池上对LLM进行微调。这从根本上受到限制,因为只有当策略rollout混合成功与失败时才会产生学习信号,导致任何固定池中有用的部分随着模型改进而迅速变得过时。为解决这一问题,我们提出前沿学习(frontier learning),一种开放式后训练方法,其中程序化生成器在线持续产生信息丰富的训练问题。它将生成器的任务特定参数视为搜索空间,并使用遗憾信号来优先排序和探索前沿难度级别,以将训练集中在模型不断演化的推理能力的边缘。在多个推理任务和模型家族中,我们的方法相对于固定池基线持续取得更高的相对增益,表明有效的后训练不仅需要选择有用的问题,还需要在能力边缘持续生成它们。

英文摘要:

Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.

↑