AI 中文总结
研究语言智能体多跳推理链变长性能下降问题,提出SLEUTH方法通过结构化认知工作记忆提升推理,经实验优势随难度增加,还发现证据充分性问题及解决办法,表明组织推理方式是扩展多跳推理关键。
AI 中文摘要
当推理链变长时,即使每个单独步骤都很简单,将推理和工具使用交织在一起的语言智能体性能也会急剧下降。我们将此归因于上下文稀释:智能体的调查状态(已确认的、怀疑的和仍需探索的)仅隐含在不断增长的上下文窗口中,早期发现被后期检索所掩盖。我们引入了SLEUTH,它通过结构化的认知工作记忆使这种状态变得明确且可操作:智能体维护基于来源的已确认事实、按证据排序的活跃假设以及直接驱动其下一步行动的未解决问题。在五个多跳基准测试和五个既定基线中,SLEUTH的优势随着难度增加而增大,从在HotPotQA上提高5分,到在4跳链上提高11分,超过了没有多次情节的Reflexion。分析剩余差距所在,我们发现了证据充分性问题:智能体常常找到答案却未能做出决断,在不必要的验证上耗尽预算。一个轻量级的决断触发器解决了这个问题,但仅当智能体已经维护结构化状态时才行:应用于非结构化智能体的相同触发器没有带来改进,这表明有组织的认知状态是有效决断的必要条件。最后,在较弱模型上强制遵守协议可在最困难的问题上恢复高达19分,这表明智能体组织推理的方式而非原始模型能力,是扩展多跳推理的关键因素。
英文摘要
Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed, what it suspects, and what it still needs) lives only implicitly in a growing context window, where early discoveries are buried under later retrievals. We introduce SLEUTH, which makes this state explicit and actionable through a structured epistemic working memory: the agent maintains Confirmed Facts grounded to sources, Active Hypotheses ranked by evidence, and Open Questions that directly drive its next action. Across five multi-hop benchmarks and five established baselines, SLEUTH's advantage grows with difficulty, from +5 points on HotpotQA to +11 on 4-hop chains, surpassing Reflexion without multiple episodes. Analyzing where the remaining gap lies, we identify the evidence sufficiency problem: agents often find the answer but fail to commit, exhausting their budget on needless verification. A lightweight commitment trigger fixes this, but only when the agent already maintains structured state: the identical trigger applied to an unstructured agent yields no improvement, isolating organized epistemic state as the necessary condition for effective commitment. Finally, enforcing protocol adherence on a weaker model recovers up to +19 points on the hardest problems, showing that how an agent organizes its reasoning, not raw model capability, is the active ingredient for scaling multi-hop reasoning.