arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09217cs.LGcs.AI

超越可解性:任务可学习性作为大语言模型强化学习后训练的静态先验

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出TrajVal估计任务可学习性,将其作为静态先验用于LLM的RL后训练任务采样,可提升数据效率并与在线调度方法互补。

中文摘要 AI 辅助

强化学习(RL)已成为从大语言模型中引出推理能力的核心后训练范式,但均匀任务采样在分配计算资源时未考虑任务对优化的响应差异。现有任务评估方法大多依赖基于快照的信号,如当前通过率或奖励,这些信号估计任务在当前策略下的可解性。然而,当前可解性相似的任务在进一步训练时的积极响应仍可能存在显著差异。我们将这一剩余维度研究为任务可学习性:在固定RL后训练机制下,任务对持续训练的预期积极响应的机制条件度量。通过分析每个任务的奖励轨迹,我们发现可学习性在独立采样的训练环境中具有可复现性,且可预测下游效用。为使该信号在训练开始前即可实用,我们提出TrajVal,一种轻量的基于探测的估计器,通过一次短探测运行和两次端点评估来近似每个任务的可学习性。TrajVal既可用作任务采样的独立静态先验,也可用作现有在线调度器的乘性先验。在多个模型规模的数学和逻辑推理基准上进行的实验表明,TrajVal相比均匀采样提高了数据效率,且与在线调度方法结合时能提供互补增益。

英文摘要

Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. Existing task-valuation methods mostly rely on snapshot-based signals such as current pass rate or reward, which estimate how solvable a task is under the current policy. However, tasks with similar current solvability can still differ substantially in how positively they respond to further training. We study this residual axis as task learnability: a regime-conditional measure of expected positive response to continued training under a fixed RL post-training regime. By analyzing per-task reward trajectories, we find that learnability is reproducible across independently sampled training contexts and predictive of downstream utility. To make this signal practical before training begins, we propose TrajVal, a lightweight probe-based estimator that approximates per-task learnability from a short probe run and two endpoint evaluations. TrajVal can be used either as a standalone static prior for task sampling or as a multiplicative prior for existing online schedulers. Experiments on mathematical and logical reasoning benchmarks across multiple model scales show that TrajVal improves data efficiency over uniform sampling and provides complementary gains when combined with online scheduling methods.

发表机构

  • Sun Yat-Sen University(中山大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑