arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05584cs.LGcs.AIcs.CL

扩展LLM推理

Expanding LLM Reasoning

Rian Atri, Evan Luo

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究LLM推理链中额外计算的最佳重启位置,提出扩展效用度量,发现学习路由优于均匀放置,但总是最后规则是强基线,并在特定模型上超越自洽前沿。

中文摘要 AI 辅助

额外的推理计算通常用于采样更多的推理链。我们研究在现有链内部,额外的延续应从何处开始。我们定义了扩展效用,即从存储的步骤重新启动链时正确性的变化,并在六个基准测试的九个模型(41个模型和基准单元)的每个合格步骤上测量它。重启位置很重要:在一组延续上选择的步骤,在不相交的集合上评分时,优于均匀放置,在5、16和38个单元的保留审计中(在新颖的五单元审计中+4.25分[+2.51,+6.63])。一个固定规则,即从最后一个合格步骤重启,即总是最后,是一个强基线:我们学习到的路由器优于均匀放置,但未检测到超过它的增益,在DeepSeek-R1-Distill-Qwen-14B/MATH-500上,总是最后在匹配的聚合生成输出上超过精确自洽前沿+0.052[+0.008,+0.098],使用四样本自洽聚合生成输出的0.774倍。交叉拟合的预言机选择仍然在声明的位置类别之外找到保留的余量,这是未来选择器的目标。最后,通过最早索引打破步骤标签平局,在每次种子诊断中,点式选择器相对于均匀放置的增益符号翻转,每次步骤有四次推出;随机平局消除了偏差。

英文摘要

Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and measure it at every eligible step for nine models on six benchmarks (41 model and benchmark cells). Restart position matters: steps selected on one set of continuations beat uniform placement when scored on disjoint ones, in held-out audits on 5, 16, and 38 cells (+4.25 points [+2.51, +6.63] in a fresh five-cell audit). A fixed rule that restarts from the last eligible steps, always-last, is a strong baseline: our learned router beats uniform placement but shows no detected gain over it, and on DeepSeek-R1-Distill-Qwen-14B/MATH-500 always-last exceeds the exact self-consistency frontier at matched aggregate generated output by +0.052 [+0.008, +0.098], using 0.774x the aggregate generated output of four-sample self-consistency. Cross-fitted oracle selection still finds held-out headroom beyond declared positional classes, a target for future selectors. Finally, breaking step-label ties by earliest index flips the sign of a pointwise selector's gain over uniform placement in every seed of a five-seed diagnostic with four rollouts per step; randomized ties remove the bias.

发表机构

  • Keiji AI
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑