arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

广泛推理,而非深度推理:将推理溢价分摊到提炼技能中

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani

arXiv 2608.07885首次发表:更新:

发表机构

Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出将推理模式的重复分摊为跨任务提炼的技能,在四个智能体基准上大幅降低代币量的同时恢复多数推理性能,部分场景下优于推理模式。

AI 中文摘要

语言模型的推理模式在多步骤智能体任务上的表现优于非推理模式,但每一轮任务的输出代币量要高出3-6倍,其中大部分代币用于重新推导同一领域各任务间共享的程序。本文表明这种重复成本可以被分摊:一个编码智能体分析训练集中一小段现有轨迹语料,并编译成一段紧凑的自然语言技能,将其注入非推理模型的系统提示中。在四个智能体基准(ALFWorld、tau²-bench电信与零售领域、SpreadsheetBench-Verified)上,该技能为GPT-5.4-mini在未见过的任务中恢复了55%-100%以上的推理差距——在四个基准中的两个上甚至直接超过了推理模式,同时输出代币量减少了2.7-6倍,且无推理代币。值得注意的是,推理轨迹并非必要条件:仅从非推理轨迹提炼出的技能,与从配对的推理/非推理语料中提炼出的技能相比仍具有竞争力,两种来源之间存在依赖领域的差异。我们通过搜索视角解释这些结果:测试时的推理是单轮任务内的深度搜索,每次部署都需重复付出成本;而语料提炼是跨任务的广泛搜索,仅需付出一次成本。两者恢复了重叠的程序知识,而利用低成本轨迹的广度通常是更优选择——部分领域(电信、SpreadsheetBench)上的剩余差距则表明,在这些领域仍需要真正的单实例深度搜索。

英文摘要

Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show this recurring cost can be amortized: a coding agent analyses a small corpus of existing trajectories from a training split and compiles a compact natural-language skill that is injected into the non-reasoning model's system prompt. Across four agentic benchmarks (ALFWorld, tau$^2$-bench telecom and retail, and SpreadsheetBench-Verified), skills recover 55%-100%+ of the reasoning gap for GPT-5.4-mini on held-out tasks -- exceeding the reasoning mode outright on two of four -- while emitting 2.7-6x fewer output tokens and zero reasoning tokens. Notably, reasoning traces are not a prerequisite: skills distilled from non-reasoning trajectories alone remain competitive with skills distilled from paired reasoning/non-reasoning corpora, with domain-dependent differences between the two sources. We interpret these results through a search lens: test-time reasoning is deep search inside a single episode, re-paid at every deployment, while corpus distillation is wide search across episodes, paid once. The two recover overlapping procedural knowledge, and width over cheap trajectories is often the better buy -- with the residual gap on some domains (telecom, SpreadsheetBench) delineating where genuinely per-instance deep search remains necessary.

CommentsCOLM 2026 Efficient Reasoning Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑