arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CIVI:用于诊断政务信息搜索智能体失败的框架

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

Dingying Liu, Yunshun Zhong, Wentao Zhang, Yiyuan Li

arXiv 2609.08094首次发表:更新:

发表机构

University of Sydney; University of Toronto; University of Waterloo; UNC-Chapel Hill(悉尼大学; 多伦多大学; 滑铁卢大学; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对政务信息搜索智能体失败问题,提出首个诊断框架CIVI及分解方法ARISE,评估十个前沿智能体均未达人类基线,72.1%失败源于检索限制。

AI 中文摘要

大型语言模型越来越多地部署在公共部门场景中,在这些场景中,不正确的指导可能导致不可逆转的伤害。我们引入了CIVI,这是第一个用于诊断政务信息中搜索智能体失败的框架。其基准实例化联合涵盖了跨国、跨辖区的政府背景(联邦、州和地方)以及来自国际采用的联合国标准的功能类别。我们评估了十个前沿搜索智能体,发现没有一个能与细心的人类基线相匹配。除了准确性之外,CIVI还衡量搜索调用率、选择性不搜索准确性以及智能体引用权威政府来源的频率。为了执行这一诊断,我们引入了ARISE,它将智能体搜索失败分解为四种互斥模式,通过来源注入消融进行隔离。ARISE将观察到的所有失败中的72.1%归因于检索受限的原因,而非模型参数化知识的差距。

英文摘要

Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause irreversible harm. We introduce CIVI, the first framework for diagnosing search agent failures in civic information. Its benchmark instantiation jointly spans cross-national, interjurisdictional government contexts (federal, state, and local) and functional categories from an internationally adopted United Nations standard. We evaluate ten frontier search agents and find that none matches an attentive human baseline. Alongside accuracy, CIVI measures search invocation rate, selective no-search accuracy, and how often agents cite authoritative government sources. To perform this diagnosis, we introduce ARISE, which decomposes agentic search failures into four mutually exclusive modes, isolated via source-injection ablation. ARISE attributes 72.1% of all observed failures to retrieval-bound causes rather than to gaps in the models' parametric knowledge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑