arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

搜索引擎从不说“不”:当检索工具拒绝时冻结智能体如何反应

Search Engines Never Say No: How Frozen Agents React When the Retrieval Tool Refuses

Ramraj Chandradevan, Sayontan Ghosh, Vinoth Selvendran

arXiv 2610.05348首次发表:更新:

AI 中文总结

本研究通过索引空洞测试平台,考察冻结搜索智能体在检索工具拒绝时的行为,发现单句拒绝能显著提升弃权率、减少错误答案,并揭示措辞影响与检测器瓶颈。

AI 中文摘要

搜索工具从不说“不”:即使索引中没有答案,它也会返回其前k个段落,因此智能体看到的是无关文本,而非未命中信号。我们转而探究当工具拒绝时,冻结的搜索智能体会作何反应。在一个索引空洞测试平台(257个NQ问题和300个HotpotQA问题,在2100万段落的BM25索引中,分别在有和没有其黄金段落的情况下运行)上,七个智能体接收到五种拒绝措辞之一。一条未事先告知的单句拒绝,使Qwen3-8B/32B在不可回答问题上的弃权(不执行)率平均从23%提升至97%,使Claude Haiku 4.5的弃权率从28%提升至57%,错误答案几乎一对一地减少,平均比系统提示指令高出51个百分点。Search-R1忽略拒绝并编造检索结果;Claude Sonnet 5.5和Opus 5.5则凭记忆作答(弃权率+2个百分点),并遵从系统提示指令(+16)。我们还发现措辞很重要:解释优于裸令牌;观察中的指令对Haiku起决定性作用;软警告毫无用处。从轻量级基于分数的预测器到LLM接地性评判器的现实触发器,远不及神谕,且都落在一条收益与信号质量曲线上,该曲线以固定误拒预算下的召回率为任何触发器定价:对于顺从的智能体,瓶颈在于工具内部的检测器,而非智能体本身,且该曲线告诉未来的检测器工作每个召回率点值多少。

英文摘要

A search tool never says no: it returns its top-k passages even when the index holds no answer, so the agent sees irrelevant text instead of a miss signal. We ask what frozen search agents do when the tool refuses instead. On an index-hole testbed (257 NQ and 300 HotpotQA questions run with and without their gold passages in a 21M-passage BM25 index), seven agents receive one of five refusal wordings. An un-announced one-sentence refusal raises abstention on unanswerable questions from 23% to 97% on average for Qwen3-8B/32B and from 28% to 57% for Claude Haiku 4.5, cutting wrong answers almost one-for-one and beating a system-prompt instruction by 51 points on average. Search-R1 ignores the refusal and fabricates retrievals; Claude Sonnet 5.5 and Opus 5.5 answer from memory (abstention +2 points) and obey a system-prompt directive instead (+16). We also found that wording matters: an explanation beats a bare token; a directive inside the observation is decisive for Haiku; a soft warning is useless. Realistic triggers, from a lightweight score-based predictor to an LLM grounding judge, fall well short of the oracle, and all land on a benefit-versus-signal-quality curve that prices any trigger by its recall at a fixed false-refusal budget: for compliant agents the bottleneck is the detector inside the tool, not the agent, and the curve tells future detector work what each point of recall is worth.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑