arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

干净分数、埋藏证据与自信错误:对前沿智能体问答的基于凭证的审计

Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA

Luis M. Sánchez

arXiv 2609.15319首次发表:更新:

发表机构

Toryx Inc.(托里克斯公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控数据室审计发现,前沿智能体在埋藏证据条件下准确率下降、成本上升,且自信错误难被校准捕捉,提出声明级凭证、条件感知评分和人类对抗性验证的审计框架。

AI 中文摘要

前沿模型在浅层文档/图表阅读任务上得分很高。在一项受控数据室审计中,将证据移至埋藏条件后,准确率下降,强制声明增加,工具调用增加,且每个正确答案的成本上升。置信度与基准校准未能完全捕捉错误答案;一个有记录的生产事故表明,捏造的结构性声明可能与准确的数值表格混杂在一起。智能体评估需要声明级凭证(陈述级来源,而非答案级分数)、条件感知评分以及人类对抗性验证——这是一种审计纪律,而非排行榜。我们测量的场景是财务尽职调查;我们下一步构建的场景是国防参谋工作,其中存在相同的埋藏证据形态。在两种场景中,模型都不是后果的当事方;签字的人才是。简而言之:在我们检查的有记录案例中,智能体可以将准确的数字与自信的捏造解释配对,因此举证责任必须从模型转移到证据链上。

英文摘要

Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced accuracy, increased forced declarations, increased tool calls, and increased cost per correct answer. Confidence and benchmark calibration did not fully capture wrong answers; a documented production incident shows fabricated structural claims can be mixed with accurate numeric tables. Agentic evaluations need claim-level receipts (statement-level provenance, not answer-level scores), condition-aware scoring, and human-adversarial verification - an auditing discipline, not a leaderboard. The setting we measure is financial due diligence; the setting we are building toward next is defense staff work, where the same buried-evidence shape appears. In both, the model is not a party to the consequences; the person who signs is. In plain terms: in the documented cases we examine, agents can pair accurate numbers with confident fabricated explanations, and the burden of proof must therefore move from the model to the evidence trail.

Comments28 pages, 9 figures. Frozen evidence archive: https://doi.org/10.5281/zenodo.22310532

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑