发表机构
Korea University; Konkuk University(高丽大学; 建国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ThinkFuse提出一种无需训练的测试时融合框架,通过比较片段级与轨迹级不确定性变化,选择性干预不稳定推理,提升小型推理模型在数学和知识密集型任务上的性能,且更高效。
AI 中文摘要
小型推理模型(SRMs)通过生成扩展的思维链轨迹,在复杂推理任务上展现出强劲性能,但一旦其推理进入错误路径,往往难以恢复。现有的测试时融合方法依赖局部融合信号来决定何时触发融合,这可能会被瞬时的不确定性波动所误导,并可能强化不稳定的推理轨迹。我们提出ThinkFuse,一种无需训练的测试时融合框架,能够选择性地干预不可靠的推理片段。ThinkFuse通过比较片段级不确定性变化与轨迹级不确定性趋势,来识别不稳定的推理点,并将辅助推理路径融合到主模型的轨迹中。大量实验表明,ThinkFuse在数学和知识密集型推理基准上优于基线方法,在模型家族组合上取得了一致的性能提升,并且在主模型较小的情况下依然保持稳健。我们的分析显示,ThinkFuse需要更少的融合触发次数并生成更少的令牌,突显了选择性触发的效率。我们的代码可在以下https URL获取。
英文摘要
Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model's trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at https://github.com/js-lee-AI/ThinkFuse.
CommentsAccepted to EMNLP 2026 Findings