arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10463cs.AIcs.IR

GRASP:用于智能检索增强生成的粒度感知搜索策略

GRASP: GRanularity-Aware Search Policy for Agentic RAG

Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan

首次发表
浏览论文内容

中文总结 AI 辅助

研究智能体检索增强生成中模型决策难题,提出GRASP强化学习框架,训练智能体协调互补检索工具,实验表明其提升检索召回率和问答性能,还展现出可解释行为,凸显协调检索信号和上下文粒度对智能体推理的关键作用。

中文摘要 AI 辅助

智能检索增强生成(Agentic RAG)通过允许语言模型迭代推理、生成搜索查询、检索证据和预测答案来扩展静态RAG。然而,模型在决定何时检索、使用词汇匹配还是语义相似性以及如何控制上下文粒度以防止无关令牌干扰智能体推理方面仍然具有挑战性。本文介绍了GRASP,这是一个强化学习框架,用于训练智能体在多步推理过程中自适应地协调互补检索工具。GRASP为智能体提供语义搜索、关键词搜索和段落阅读操作,使其能够仅在需要时检索句子级证据并扩展进一步的上下文。我们使用一种联合考虑答案准确性、基于事实的阅读、互补搜索和轮次效率的奖励来训练策略。在多跳推理基准上的实验表明,与单步检索、基于提示的智能体RAG和基于强化学习的检索基线相比,GRASP提高了检索召回率和下游问答性能。定性和消融分析表明,学习到的策略发展出了可解释的略读和扫描行为:它使用语义搜索进行广泛探索,段落阅读进行局部验证,关键词搜索用于特定实体的证据。这些结果表明,学习协调检索信号和上下文粒度对于智能体的正确推理至关重要。

英文摘要

Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matching or semantic similarity, and how to control context granularity to prevent irrelevant tokens from interfering with agent reasoning. In this paper, we introduce GRASP, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning. GRASP provides the agent with semantic search, keyword search, and paragraph-reading actions, enabling it to retrieve sentence-level evidence and expand further context only when needed. We train the policy with a reward that jointly accounts for answer accuracy, grounded reading, complementary search, and turn efficiency. Experiments on multi-hop reasoning benchmarks show that GRASP improves both retrieval recall and downstream question answering performance compared with single-step retrieval, prompting-based agentic RAG, and RL-based retrieval baselines. Qualitative and ablation analyses show that the learned policy develops interpretable skimming and scanning behavior: it uses semantic search for broad exploration, paragraph reading for local verification, and keyword search for entity-specific evidence. These results suggest that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.

发表机构

  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Adobe Research(Adobe 研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑