arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01867cs.CL

CRISP:用于训练高效深度搜索智能体的关键步骤感知

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Haosi Mo, Zihao Yan, Ruiqing Zhang, Zhongli Li, Hexuan Deng, Xuebo Liu, Min Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对深度搜索智能体效率低的问题,提出CRISP框架,通过区分关键与冗余交互优化奖励,在保持准确率的同时,使BrowseComp和HLE-Verified上的交互回合分别减少15.1%和33.2%。

中文摘要 AI 辅助

大型语言模型(LLMs)正被越来越多地扩展为深度搜索智能体,这些智能体通过与外部搜索和浏览工具的多步交互来解决复杂问题。然而,现有的智能体通常会产生大量的计算和交互成本,生成包含冗余查询、低效探索和无关观察的冗长轨迹。现有的面向效率的方法通常鼓励智能体减少工具使用频率,但对所有工具交互一视同仁可能会抑制收集必要证据的步骤。在本文中,我们提出CRISP,一种通过关键步骤感知来训练高效深度搜索智能体的框架。与之前对工具使用进行统一惩罚的效率方法不同,CRISP区分收集必要证据的交互和冗余交互,并调整训练奖励以保留前者同时修剪后者,在不牺牲正确答案所需证据的情况下提高效率。具体而言,CRISP首先通过反向证据归纳构建关键步骤标签:从最终答案开始,一个强模型向后遍历完整的搜索轨迹并判断每个工具交互步骤是否提供或保留了最终答案的证据。然后我们将这些逐步骤判断提炼成一个更小的关键步骤识别器,支持单次通过的全轨迹分析。在策略优化期间,仅对成功的 rollout 应用感知效率的奖励。在BrowseComp和HLE-Verified上的实验表明,CRISP保持了有竞争力的最终答案准确率,同时分别将平均交互回合减少了15.1%和33.2%,展现出交互效率的显著提升。

英文摘要

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools. However, existing agents often incur substantial computational and interaction costs, generating lengthy trajectories that contain redundant queries, inefficient exploration, and irrelevant observations. Existing efficiency-oriented methods usually encourage agents to use tools less frequently, but treating all tool interactions uniformly may also suppress steps that gather necessary evidence. In this paper, we propose CRISP, a framework for training efficient deep search agents through critical step perception. Unlike prior efficiency methods that uniformly penalize tool use, CRISP distinguishes interactions that gather necessary evidence from redundant ones and shapes the training reward to preserve the former while pruning the latter, improving efficiency without sacrificing the evidence needed for correct answers. Specifically, CRISP first constructs critical-step labels with Backward Evidence Induction: starting from the final answer, a strong model traverses a completed search trajectory backward and judges whether each tool-interaction step provides or preserves evidence for the final answer. We then distill these step-wise judgments into a smaller critical-step recognizer, enabling full-trajectory analysis in a single pass. During policy optimization, an efficiency-aware reward is applied only to successful rollouts. Experiments on BrowseComp and HLE-Verified show that CRISP maintains competitive final-answer accuracy while reducing average interaction turns by 15.1% and 33.2%, respectively, demonstrating substantial improvements in interaction efficiency.

补充信息

↑