arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26070cs.CLcs.AIcs.LG

用于高效测试时缩放的前缀滑动

Prefix Sliding for efficient test-time scaling

Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, L… 展开作者

Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt, Percy Liang, Jason Wei, Andrew Y. Ng, Luke Zettlemoyer, Yejin Choi, Mike Lewis

首次发表
浏览论文内容

中文总结 AI 辅助

针对测试时推理内存成本过高问题,提出Prefix Sliding方法,通过丢弃非关键标记限制内存,无需训练可提速3倍,结合强化学习训练后性能更优且优于其他基线

中文摘要 AI 辅助

测试时缩放利用额外的测试时计算提升性能,例如让语言模型在解决问题时进行更长时间的推理。由于模型通过全注意力机制将整个推理轨迹保留在内存中,需要长时间思考的困难任务会因计算成本过高而难以承受。然而,我们发现随着模型持续推理,大多数中间推理标记会失去重要性,这引发了是否值得为保留这些标记付出成本的疑问。基于这一见解,我们提出Prefix Sliding(前缀滑动),其在推理过程中会丢弃不属于前缀或最近数千个标记窗口的标记。前缀包含模型可用的关键指令和工具,而最近的标记是模型当前正在进行的推理,无论模型推理多长时间,这都能限制总内存需求,实现高效的长程测试时缩放。无需训练,Prefix Sliding可使现有模型速度提升3倍,同时保持性能;使用强化学习结合Prefix Sliding进行训练,可通过扩展至十万个标记以上的推理轨迹实现更好的性能。消融实验表明,Prefix Sliding的性能优于中间标记摘要或普通滑动窗口。我们的代码位于this https URL

英文摘要

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. However, we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens. The prefix has key instructions and tools available to the model, while the most recent tokens are the current reasoning the model is working on. This caps the total memory requirement regardless of how long the model reasons, allowing for efficient long-horizon test-time scaling. Without training, Prefix Sliding can make existing models 3x faster while maintaining performance. Training with Prefix Sliding using reinforcement learning can achieve better performance by enabling scaling to reasoning traces beyond a hundred thousand tokens. Ablations show Prefix Sliding outperforms summarizing intermediate tokens or vanilla sliding window. Our code is at https://github.com/Muennighoff/prefix-sliding

发表机构

  • Stanford University(斯坦福大学)
  • University of California at Santa Barbara(加州大学圣巴巴拉分校)
  • Prime Intellect
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑