arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13237cs.IRcs.CL

多轮检索增强生成(RAG)应何时停止?Search-R1中的结构化停止判断与检索减少

When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1

Weimeng Luo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对多轮RAG的停止判断问题,将S2G-RAG的结构化判断适配到Search-R1,训练Qwen3.5-2B判断器,可减少约3.70%检索调用且仅小幅降低答案准确率。

中文摘要 AI 辅助

多轮检索增强生成(RAG)必须在证据积累到一定程度时决定何时停止检索。由于部署的策略由每条轨迹上的第一个STOP决定,这是一个序列选择问题,而非独立的状态分类任务。我们将S2G-RAG的结构化充分性与差距判断适配到冻结的Search-R1流程中,并基于900个不重叠HotpotQA问题的3009个状态训练了一个Qwen3.5-2B判断器。Search-R1的推理器、检索器、语料库、提示词和检索预算保持不变,而判断器检查点和停止阈值在分组验证集上选择并冻结,用于确认性评估。在确认性测试集上,与原生Search-R1相比,该策略减少了77次检索调用(降幅3.70%),同时官方精确匹配率下降了0.625个百分点。因此,训练后的S2G风格结构化判断器在减少检索的同时,大致保留了答案准确性。该结果并不意味着准确性不变或提升、安全停止或总推理成本降低。

英文摘要

Multi-round retrieval-augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem rather than an independent state-classification task. We adapt S2G-RAG's structured sufficiency-and-gap judgment to a frozen Search-R1 pipeline and train a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions. Search-R1's reasoner, retriever, corpus, prompt, and search budget remain unchanged, while the judge checkpoint and stopping threshold are selected on grouped validation and frozen before confirmatory evaluation. On the confirmatory test set, the resulting policy reduces retrieval calls by 77 (3.70\%) relative to Native Search-R1, while Official Exact Match decreases by 0.625 percentage points. Thus, the trained S2G-style structured judge reduces retrieval while broadly preserving answer accuracy. The result does not imply unchanged or improved accuracy, safe stopping, or lower total inference cost.

补充信息

↑