AI 中文总结
该研究针对深度研究智能体的推理停滞问题,提出检索感知智能体控制器(RAAC),经实验验证可减少搜索调用次数、提升智能体的召回率与准确率。
AI 中文摘要
在本文中,我们分析了多种深度研究智能体(DRA)的推理轨迹,发现现有智能体常存在推理停滞问题:多数迭代对最终性能几乎无改进,且智能体缺乏对自身轨迹的感知,无法有效调整搜索策略或确定终止时机。为解决该问题,我们引入一组无监督信号及检索感知智能体控制器(RAAC),协助智能体在研究各阶段选择最优动作。RAAC融入搜索新颖性、信息覆盖度等关键信息检索原则,形成更高效的推理轨迹,在提升整体性能的同时减少不必要迭代,进而降低成本与延迟。在BrowseComp-Plus数据集及多种DRA上,加入RAAC后搜索调用次数平均减少14次,显著提升了性能最优的DRA的召回率与准确率,准确率提升最高达10%(平均3%)。
英文摘要
In this paper, we analyze the reasoning trajectories of a variety of DRAs and show that existing agents often suffer from reasoning stagnation: the majority of iterations contribute little or no improvement to final performance, while agents lack awareness of their trajectories and are therefore ineffective at adapting their search strategies or determining when to terminate. To address this issue, we introduce a set of unsupervised signals and a Retrieval-Aware Agent Controller (RAAC), which assists the agent in selecting optimal actions at each stage of the research process. RAAC incorporates key information retrieval principles, namely search novelty and information coverage, resulting in more effective reasoning trajectories that improve overall performance while reducing unnecessary iterations, and consequently cost and latency. Specifically on BrowseComp-Plus and across a large set of DRAs, adding RAAC reduces the number of search calls by an average of 14, significantly improves the best-performing DRA on recall and accuracy, and achieves an accuracy gain of up to 10% (3% on average).