arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓解长度扩展税:在线蒸馏方法

Mitigating the Length-Scaling Tax with Online Distillation

Xu Wan, Wenyue Xu, Shengjie Zhao, Mingyang Sun

arXiv 2609.38854首次发表:更新:

发表机构

ByteDance Seed; Tongji University; Peking University(字节跳动Seed; 同济大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出长度自蒸馏(LSD)方法,通过在线策略蒸馏缓解RL后训练中的长度扩展税,在保持性能的同时显著减少简单查询的回答长度,实验显示LST从19.0%降至-3.7%(单轮)和从31.4%降至13.7%(多轮)。

AI 中文摘要

在强化学习(RL)后训练期间的长度扩展通常被视为推理能力提升的标志,尤其是在处理难题时,但这也可能导致对已解决问题产生不必要的冗长回答。我们将这种副作用量化为长度扩展税(LST):在已解决的查询上,回答长度过度增加而准确率没有相应提升。为缓解LST,我们提出了长度自蒸馏(LSD)方法,该方法将已解决的提示路由到在线策略蒸馏,而对未解决的提示保留原始RL目标。LSD使用在线策略的指数移动平均作为其教师模型,无需外部模型。我们发现,在多种变体下,LSD在性能上与RL相当或更优,同时显著抑制了简单查询上回答长度的增长。LSD将单轮推理中的LST从19.0%降至-3.7%,在多轮智能体任务中从31.4%降至13.7%,表明LSD在RL后训练期间能有效保持简单查询上的简洁回答模式,同时支持对困难查询的高效探索。

英文摘要

Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑