arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05139cs.CLcs.LG

面向技能原生的大语言模型:用于基准测试和训练长程推理的技能熵

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

  • Princeton University(普林斯顿大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Toronto(多伦多大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Stanford University(斯坦福大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora

AI总结:

该研究针对现有基准无法评估LLM跨技能长程推理能力的问题,提出Skill Entropy(技能熵)并构建Skill²-Bench基准,还开发Skill-Entropy RL训练框架,显著提升了Qwen3模型在该基准上的表现。

AI中文摘要:

近期大语言模型(LLM)的长程推理要求模型在推理链中切换不同技能,例如先进行数学推导,再利用结果规划日程。我们将这类问题称为跨技能长程任务:即多步骤任务,其各步骤需要不同的推理技能且依赖于早期输出。现有基准通常评估单个技能,缺乏衡量模型在技能间切换能力的有效方法。我们从评估和训练两方面解决这一缺口。我们提出Skill Entropy(技能熵),用于衡量从一个技能切换到另一个技能的难度。随后我们构建了Skill²-Bench,这是一个跨技能长程任务基准,涵盖9个可验证且开放的领域,共558项技能。每个任务都被赋予任务级技能熵分数,并分为三个难度级别。我们在Skill²-Bench上评估了8个前沿模型和4个开源模型,发现存在技能切换差距:在更高熵的任务上,模型准确率会下降。接着我们将技能熵从基准尺度转化为训练信号,提出Skill-Entropy RL,这是一个强化学习(RL)框架,其中模型不仅预测每一步的答案,还预测用于生成该答案的技能。奖励函数结合了步骤级正确性与技能熵奖励,技能熵奖励用于衡量模型预测的技能序列与黄金技能序列的一致性。在Qwen3-4B-Instruct和Qwen3-1.7B上,Skill-Entropy RL将Skill²-Bench的分数分别从34.4%提升至68.4%、从14.6%提升至40.1%,优于竞争基准。该流程还可应用于现成训练数据如OpenR1-Math,表明技能熵是一种可复用的训练信号。代码可从该https URL获取。

英文摘要:

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill^2-Bench, a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open-source models on Skill^2-Bench reveals a skill-switching gap: accuracy decreases on higher-entropy tasks. We then turn skill entropy from a benchmark scale into a training signal. We propose Skill-Entropy RL, an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step-level correctness with a skill-entropy reward that measures the alignment between the model-predicted skill sequence and the gold skill sequence. On Qwen3-4B-Instruct and Qwen3-1.7B, Skill-Entropy RL improves the Skill^2-Bench score from 34.4% to 68.4% and from 14.6% to 40.1%, respectively, outperforming competitive baselines. The same pipeline can be applied to off-the-shelf training data such as OpenR1-Math, indicating that skill entropy is a reusable training signal. Code available at: https://github.com/Gen-Verse/Skill-Entropy-RL

补充信息

相关深度报道

↑