arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

搜索塑造结论:审计深度研究智能体中的证据选择偏差

Search Shapes Conclusions: Auditing Evidence Selection Bias in Deep Research Agents

Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Lulu Wang, Ziming Yu, Junxi Yin

arXiv 2609.39026首次发表:更新:

发表机构

Beijing Normal University; Ke Holdings(北京师范大学; 贝壳控股有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出因果证据选择校正(CESS)方法,用于审计深度研究智能体在自适应搜索中的证据选择偏差,通过校正候选池平均证据方向,显著降低估计误差并区分策略干预效应。

AI 中文摘要

深度研究智能体将证据综合成带引用的报告,然而一份引用良好的报告仍可能得出误导性结论。引用正确性检查的是所引来源是否支持个别论断,它并不能表明自适应搜索是否暴露了所有可供评估文档(我们称之为候选池)的代表性视图。早期发现会重定向后续查询、文档选择和停止条件,因此智能体阅读的文档构成一个选择性样本。现有评估很少考虑这种选择。我们将该问题形式化为自适应证据采样,并引入因果证据选择校正(CESS)。CESS 预测每个候选文档的证据方向,并利用选择每个文档及到达每个搜索轮次的记录概率来校正候选池平均值。收缩(shrinkage)稳定短搜索,而当某些文档无法采样时,用区间替代点估计。我们还证明,估计公共池的平均证据方向不同于衡量搜索策略变化如何改变所读证据;后者需要干预。在 MS2 系统综述基准的问题上,相对于阅读文档证据分数的平均值,CESS 将候选池平均值的平均绝对误差降低了 9.2%,并将估计值在相反文档排序下的变化降低了 39.4%。在来自公开 Open Deep Research 智能体的轨迹中,相应降幅分别达到 60.1% 和 87.2%。在配对干预下的额外 4,800 条轨迹证实,校正池估计与衡量策略效应是不同任务。因此,CESS 审计报告背后的证据方向是否反映可供评估的文档,而单独的干预分析则衡量搜索决策的效应。

英文摘要

Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation, which we call the candidate pool. Early findings redirect later queries, document choices, and stopping, so the documents an agent reads form a selective sample. Existing evaluations rarely account for this selection. We formulate the problem as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts each candidate document's evidence direction and corrects the candidate-pool average using the logged probabilities of selecting each document and reaching each search round. Shrinkage stabilizes short searches, while intervals replace point estimates when some documents cannot be sampled. We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read. The latter requires intervention. On questions from the MS2 systematic-review benchmark, CESS reduces mean absolute error against the candidate-pool average by $9.2\%$ and reduces the estimate's change under opposing document rankings by $39.4\%$ relative to averaging the evidence scores of documents read. Across trajectories from a public Open Deep Research agent, the corresponding reductions reach $60.1\%$ and $87.2\%$. A further 4,800 trajectories under paired interventions confirm that correcting a pool estimate and measuring a policy effect are different tasks. CESS therefore audits whether the evidence direction underlying a report reflects the documents available for evaluation, while a separate intervention analysis measures the effect of search decisions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑