arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DRNOISE:在误导性证据环境中对深度研究代理进行基准测试

DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments

Jun Nie, Zhiqin Yang, Zhenheng Tang, Yonggang Zhang, Xiaowen Chu, Xinmei Tian, Bo Han

arXiv 2607.17291首次发表:更新:

发表机构

Hong Kong Baptist University; University of Science and Technology of China; The Hong Kong University of Science and Technology(香港浸会大学; 中国科学技术大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在误导性证据环境下深度研究代理的表现,引入DRNOISE基准,每个任务含正确答案及冲突文档,涵盖多类证据操作。测试发现代理存在验证惰性,通用提示可缩小差距,强调可靠深度研究需积极协调主张与证据。

AI 中文摘要

深度研究代理越来越多地在开放网络上运行,相关记录与冗余摘要、过时报告和误导性文档共存。现有评估对于当一份看似普通的虚假文档被故意植入可搜索环境并提供与冲突答案的直接捷径时,代理是否能保持合理的证据标准了解有限。我们引入了DRNOISE,一个用于在误导性证据下答案恢复的100任务基准。每个任务都有一个由两个相互佐证的间接记录链支持的唯一正确答案;配对的噪声条件添加了一个直接给出冲突答案的似是而非的文档。该基准涵盖十个证据操作类别。在具有强大的干净任务性能的代理中,这种单一干预导致准确率下降66 - 88个百分点。跟踪分析确定验证惰性是主要的失败模式:代理经常检索到真实记录,但在完成和协调证据链之前就停止了,而是听从类似答案的文档。通用验证提示缩小了但并未消除这一差距。该设置与开放网络部署特别相关,在开放网络中似是而非的虚假信息通过看似普通的页面而非明确攻击出现。因此,可靠的深度研究不仅需要检索和引用;还需要将直接主张与记录级证据进行积极协调。

英文摘要

Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-looking false document is deliberately seeded into a searchable environment and offers a direct shortcut to a conflicting answer. We introduce DRNOISE, a 100-task benchmark for answer recovery under misleading evidence. Each task has a unique gold answer supported by two corroborating indirect record chains; the paired noisy condition adds one plausible document that states a conflicting answer directly. The benchmark spans ten families of evidence operations. Across agents with strong clean-task performance, this single intervention causes 66-88 percentage-point accuracy drops. Trace analyses identify verification inertia as the dominant failure mode: agents often retrieve truthful records but stop before completing and reconciling the evidence chain, instead deferring to the answer-like document. Generic verification prompts reduce but do not close this gap. The setting is especially relevant to open-web deployment, where plausible falsehoods arrive through ordinary-looking pages rather than explicit attacks. Reliable deep research therefore requires more than retrieval and citation; it requires active reconciliation of direct claims with record-level evidence.

Comments16 pages, 2 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑