arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34327cs.AIcs.CL

知道何时思考不够:教小型推理模型超越其参数知识进行推理

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang

AI总结:

本文提出FlyBy框架,教小型推理模型在知识瓶颈时选择性查询更强模型,以低成本超越大模型,实现更优推理性能。

AI中文摘要:

扩展测试时计算是提升语言模型推理能力的一种强大方式,尤其对于服务成本低廉的小型推理模型(sRMs)而言颇具吸引力。然而,额外的思考是否总是正确的操作?通过在两个模型家族和多个规模上对中间推理状态进行干预,我们发现自我精炼主要将概率质量集中到当前状态已可达的解决方案上,而非使新方案变得可达。这些干预揭示了两种失败模式:执行瓶颈,即正确路径可达且反思能够恢复它;以及知识瓶颈,即相关外部信息使其可达。受此区别启发,我们引入了FlyBy,一个选择性查询框架,并训练了4B和8B变体,使其先推理,诊断仍未解决的问题,并在知识瓶颈处查询参数知识超出自身的更强模型。监督微调引导了多深度查询动作,而成本感知的强化学习则校准是否查询、查询什么以及花费多少。在六个基准的1,158个难题上,FlyBy-4B实现了45.96%的pass@8,以2.7倍更低的服务成本超越了Qwen3-14B(41.64%),同时在pass@1上也超过了Qwen3-8B(16.85%对15.31%)。扩展到FlyBy-8B进一步将pass@8提升至51.81%。

英文摘要:

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.

补充信息

↑