arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22602cs.AI

DeepLook:通过前瞻进行更深入思考

DeepLook: Deeper Thinking with Lookahead

  • Technical University of Munich(慕尼黑工业大学)
  • LMU Munich(慕尼黑大学)
  • MCML, LMU, MemAgents Lab(慕尼黑大学机器学习中心、记忆代理实验室)

机构由 AI 辅助整理,请以论文原文为准。

Tingxin Yang, Zefeng Wang, Mengyue Wang, Xingcheng Zhou, Yunpu Ma

AI总结:

研究旨在改进大语言模型推理,提出无需训练的DeepLook框架,通过将前瞻计算集中在不确定性瓶颈处,聚合令牌级置信度为段级信号并排序投票。在四个数学基准测试中,该方法改变准确率与令牌成本的帕累托前沿,提升准确率并减少令牌生成。

AI中文摘要:

推理时缩放已成为改进大语言模型推理的强大范式,在困难推理任务上往往比仅参数缩放带来更大提升。但现有方法在推理轨迹内计算分配效率低。基于推理失败常在错误答案明确前就出现早期不确定性的观察,我们引入了DeepLook,这是一个无需训练的监控与干预解码框架,将前瞻计算集中在不确定性瓶颈处。它将令牌级置信度聚合为段级信号,在置信度相对于近期历史下降时触发,并通过固定视野前瞻探索候选延续。分支按平均前瞻置信度(ALC)排序,然后通过投票进行修剪和聚合。在四个跨模型的竞赛式数学基准测试中,DeepLook改变了准确率与令牌成本的帕累托前沿:在16种设置中的11种情况下提高了准确率,同时平均减少了87.3%的数据集级令牌生成,如使用Qwen3 - 32B在AIME25上提高了3.1,使用GPT - OSS - 20B在BRUMO25上提高了8.8。结果表明,有选择的、具有未来意识的干预比统一缩放完整推理轨迹能产生更强的准确率与成本权衡。

英文摘要:

Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter scaling alone. However, existing approaches remain inefficient in how compute is allocated within a reasoning trace. Motivated by the observation that reasoning failures often exhibit an early onset of uncertainty before a wrong answer become explicit, we introduce DeepLook, a training-free monitor-and-intervene decoding framework that concentrates lookahead compute at uncertainty bottlenecks. DeepLook aggregates token-level confidence into segment-level signals, triggers when confidence drops relative to recent history, and explores candidate continuations with fixed-horizon lookahead. Branches are ranked by Average Lookahead Confidence (ALC), the average segment-level confidence over rollout continuations, then pruned and aggregated through voting. On four competition-style mathematics benchmarks across DeepSeek-R1-8B, Qwen3-32B, GPT-OSS-20B, and GPT-OSS-120B, DeepLook shifts the accuracy--token-cost Pareto frontier: it improves accuracy over DeepConf-low in 11 of 16 settings while reducing dataset-level token generation by 87.3% on average, including gains of +3.1 on AIME25 with Qwen3-32B and +8.8 on BRUMO25 with GPT-OSS-20B. These results show that selective, future-aware intervention yields substantially stronger accuracy--cost trade-offs than uniformly scaling complete reasoning trajectories. Code is available here.

↑