发表机构
Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型推理模型多样本推理成本过高的问题,提出FoT方法,通过识别犹豫信号剪枝无产出轨迹,在保持准确率的同时大幅降低推理成本且可跨架构和任务迁移。
AI 中文摘要
大型推理模型(LRM)在同一问题的重复查询中会产生多样且有时不一致的答案,因此多样本推理是可靠部署的前提。k次展开的多数投票是该场景下的标准解决方案和事实上的准确率目标,但在LRM所需规模下其成本高得令人望而却步。我们提出思维漏斗(Funnel of Thoughts, FoT),这是一种推理时方法,可在保留完整32条轨迹投票准确率的同时,将注意力浮点运算量(FLOPs)减半,全模型推理成本降低28.8%。在来自6个LRM的11.5万条推理轨迹中,我们发现无产出轨迹常通过“等等”“其实”“或许”等重复犹豫标记显现;这些轨迹更难得出正确答案,消耗不成比例的注意力FLOPs,最坏情况下会退化为无答案循环。FoT基于这种无需训练的词汇信号,识别捕捉这些病态模式的词汇,并在完成前剪枝受影响的轨迹,无需额外模型推理即可将在线生成注意力FLOPs降低56.1%, wall time减少37.6%;该信号无需微调即可跨保留架构和域外任务迁移。
英文摘要
Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de facto accuracy target for this regime, but it is prohibitively expensive at the scale LRMs require. We introduce Funnel of Thoughts (FoT), an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost. Across 115K reasoning trajectories from six LRMs, we find that unproductive trajectories often reveal themselves through repeated hesitation markers such as "Wait", "Actually", and "perhaps." These trajectories are less likely to reach the correct answer and consume disproportionate attention FLOPs, degenerating into no-answer loops in the worst case. Built on this training-free lexical signal, FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1% and wall time by 37.6% without any additional model inference; the same signal transfers without retuning across held-out architectures and out-of-domain tasks.
Comments20 pages, 8 figures