AI 中文总结
本研究通过索引空洞测试平台,考察冻结搜索智能体在检索工具拒绝时的行为,发现单句拒绝能显著提升弃权率、减少错误答案,并揭示措辞影响与检测器瓶颈。
AI 中文摘要
搜索工具从不说“不”:即使索引中没有答案,它也会返回其前k个段落,因此智能体看到的是无关文本,而非未命中信号。我们转而探究当工具拒绝时,冻结的搜索智能体会作何反应。在一个索引空洞测试平台(257个NQ问题和300个HotpotQA问题,在2100万段落的BM25索引中,分别在有和没有其黄金段落的情况下运行)上,七个智能体接收到五种拒绝措辞之一。一条未事先告知的单句拒绝,使Qwen3-8B/32B在不可回答问题上的弃权(不执行)率平均从23%提升至97%,使Claude Haiku 4.5的弃权率从28%提升至57%,错误答案几乎一对一地减少,平均比系统提示指令高出51个百分点。Search-R1忽略拒绝并编造检索结果;Claude Sonnet 5.5和Opus 5.5则凭记忆作答(弃权率+2个百分点),并遵从系统提示指令(+16)。我们还发现措辞很重要:解释优于裸令牌;观察中的指令对Haiku起决定性作用;软警告毫无用处。从轻量级基于分数的预测器到LLM接地性评判器的现实触发器,远不及神谕,且都落在一条收益与信号质量曲线上,该曲线以固定误拒预算下的召回率为任何触发器定价:对于顺从的智能体,瓶颈在于工具内部的检测器,而非智能体本身,且该曲线告诉未来的检测器工作每个召回率点值多少。
英文摘要
A search tool never says no: it returns its top-k passages even when the index holds no answer, so the agent sees irrelevant text instead of a miss signal. We ask what frozen search agents do when the tool refuses instead. On an index-hole testbed (257 NQ and 300 HotpotQA questions run with and without their gold passages in a 21M-passage BM25 index), seven agents receive one of five refusal wordings. An un-announced one-sentence refusal raises abstention on unanswerable questions from 23% to 97% on average for Qwen3-8B/32B and from 28% to 57% for Claude Haiku 4.5, cutting wrong answers almost one-for-one and beating a system-prompt instruction by 51 points on average. Search-R1 ignores the refusal and fabricates retrievals; Claude Sonnet 5.5 and Opus 5.5 answer from memory (abstention +2 points) and obey a system-prompt directive instead (+16). We also found that wording matters: an explanation beats a bare token; a directive inside the observation is decisive for Haiku; a soft warning is useless. Realistic triggers, from a lightweight score-based predictor to an LLM grounding judge, fall well short of the oracle, and all land on a benefit-versus-signal-quality curve that prices any trigger by its recall at a fixed false-refusal budget: for compliant agents the bottleneck is the detector inside the tool, not the agent, and the curve tells future detector work what each point of recall is worth.