科学搜索代理失败之处:暴露与检查尝试的决策检查点审计
Where Scientific Search Agents Fail: Decision-Checkpoint Auditing of Exposure and Inspection Attempts
浏览论文内容
中文总结 AI 辅助
针对科学搜索代理,提出决策检查点审计方法,通过记录推理中的观察与工具操作,区分目标暴露和检查失败,在540个问题上验证了关键词搜索与先读后搜的差异。
中文摘要 AI 辅助
最终答案的准确性无法揭示科学搜索代理是未能遇到目标论文、尝试检查该论文,还是在检查后返回了可接受的答案。我们引入了决策检查点,在推理过程中无需基准标签即可记录观察结果和工具操作,然后将目标身份与评估者标签相结合,从记录的事件中分配结果类别。在固定且目标增强的环境中,针对540个可回答的AutoResearchBench深度问题,在五种条件下,关键词搜索的准确率为24.6%,而原始搜索的准确率为17.8%。关键词条件下,既无目标暴露也无检查的错误答案较少,但目标暴露后未被检查的错误答案较多。与关键词搜索相比,先读后搜条件下的记录证据搜索调用次数多27.4%。在先读后搜条件下,目标检查尝试发生在199个问题上,而在关键词搜索条件下为191个;两种条件下的准确率均为24.6%。检查点协议使这些问题级别的差异变得明确,将目标暴露和检查与总体准确率和工具使用总量区分开来。
英文摘要
Final-answer accuracy does not reveal whether a scientific-search agent failed to encounter a target paper, attempt to inspect it, or return an accepted answer after inspection. We introduce decision checkpoints that record observations and tool actions without benchmark labels during inference, then join target identities and evaluator labels to assign outcome categories from recorded events. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieves 24.6\% accuracy, compared with 17.8\% for raw search. The keyword condition has fewer incorrect answers with neither target exposure nor inspection, but more incorrect answers after the target is exposed and left uninspected. Compared with keyword search, read-first has 27.4\% more recorded evidence-search calls. Target inspection attempts occur on 199 questions under read-first and 191 under keyword search; both conditions achieve 24.6\% accuracy. The checkpoint protocol makes these question-level differences explicit, distinguishing target exposure and inspection from aggregate accuracy and total tool use.
发表机构
- Institute of Science Tokyo(东京科学大学)
- The University of Tokyo(东京大学)
- Hiroshima University(广岛大学)
机构由 AI 辅助整理,请以论文原文为准。