arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

桥梁证据:静态检索效用无法预测多步智能体搜索中的因果效用

Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search

Debayan Mukhopadhyay, Utshab Kumar Ghosh, Shubham Chatterjee

arXiv 2607.15253首次发表:更新:

发表机构

University of Calcutta; Missouri University of Science and Technology(卡塔克大学; 密苏里科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究智能体检索中静态检索效用与因果效用的差异,通过在HotpotQA上用ReAct风格智能体实验,比较原始与反事实运行得CTU分数,发现两者近乎独立,约三分之一文档为桥梁文档,还确定了相关机制,指出优化静态相关性无法带来因果有用性。

AI 中文摘要

检索系统在静态有用性概念下进行训练和评估:将一份文档和一个问题交给一个阅读器模型,查看答案是否改进,然后据此对文档评分。当单独阅读一份文档时,这个概念是成立的。但当语言模型作为搜索智能体,进行多次查询并逐轮推理时,这个概念就不成立了,因为一份文档的重要性在于它让智能体接下来做什么,而不是它对当前问题的表述。我们测量了这种差距。通过在HotpotQA上使用ReAct风格的智能体,我们重放了1000个开发问题,并在智能体读取的每个文档被删除后,从该点重新运行轨迹的其余部分。将原始运行与反事实运行进行比较,从三个增量得出反事实轨迹效用(CTU)分数:最终答案质量、下一个查询检索质量和轮数。在23322个文档观察中,将CTU与静态RAG效用(SRU)进行交叉分析,两者在统计上几乎是独立的(斯皮尔曼相关系数=-0.026)。大约三分之一的智能体读取的文档在对静态阅读器看似无用时却具有因果承载作用;我们称这些为桥梁文档。当将基于阅读器的轴换成BM25和交叉编码器代理时,这种模式仍然存在,在均匀分布的轴上给出了27.2%的桥梁单元。第二个实验确定了机制。使用先前工作中的可观察实体相关性(OER)度量,区分相关候选者和非相关候选者的实体在智能体的下一个查询中出现的频率比仅在非相关文档中发现的实体高4.02倍(6.1%对1.5%,n=227139)。桥梁文档通过为智能体提供一个引导搜索的判别实体来发挥作用。在智能体检索中,静态相关性和因果有用性是不同的量,优化前者并不能带来后者。

英文摘要

Retrieval systems are trained and evaluated on a static idea of usefulness: hand a document and a question to a reader model, see whether the answer improves, and score the document accordingly. The idea holds up when a document is read on its own. It breaks when a language model works as a search agent, issuing several queries and reasoning across turns, because a document can matter for what it lets the agent do next rather than for what it says about the current question. We measure that gap rather than argue it. Using a ReAct style agent over HotpotQA, we replay 1000 development questions and, for every document the agent read, delete it and re-run the rest of the trajectory from that point. Comparing the original run against its counterfactual gives a Counterfactual Trajectory Utility (CTU) score from three deltas: final answer quality, next query retrieval quality, and turn count. Crossing CTU against Static RAG Utility (SRU) over 23,322 document observations, the two are close to statistically independent (Spearman rho = -0.026). Roughly a third of the documents the agent reads are causally load bearing while looking useless to a static reader; we call these bridge documents. The pattern survives when the reader based axis is swapped for a BM25 and cross encoder proxy, giving a bridge cell of 27.2% on an evenly spread axis. A second experiment pins down the mechanism. Using the Observable Entity Relevance (OER) measure from prior work, entities that discriminate relevant from non-relevant candidates appear in the agent's next query 4.02 times more often than entities found only in non-relevant documents (6.1% vs 1.5%, n = 227,139). A bridge document earns its keep by handing the agent a discriminative entity that redirects the search. Static relevance and causal usefulness are different quantities in agentic retrieval, and optimizing the first does not deliver the second.

CommentsPreprint; extended version in preparation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑