arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05490cs.AIcs.LGstat.ML

自主分析智能体的创新残差审计:定位、检测限、误差控制与可识别性

Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability

Ahmed Hassoon, Mark Dredze

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对自主分析智能体的创新残差审计,分析了错误定位、分数选择、误差控制及审计极限,指出表示维度是关键约束。

中文摘要 AI 辅助

自主智能体如今可执行完整的数据分析任务,几乎无需逐步监督就能完成队列选择、表连接和模型拟合等操作。当此类分析出错时,必须确定是哪一步操作导致了错误。近期有一种方法无需任何标注错误,而是从已知正确的分析中学习,并标记出偏离该模型预测的操作,但这类审计的可靠性尚未得到研究。本文对该问题进行了分析:分数的选择决定了错误能否被定位。若根据操作相对于其直接前序操作的意外程度对每个操作打分,那么仅继承早期错误的操作会与正确操作无法区分,因此一个错误会产生一个标记;而根据更长的预期分析重构计算分数,则会将单个错误扩散到多个操作中。本文量化了这种扩散程度,以及当错误是逐渐累积而非一次性出现时如何选择比较长度。随后,本文给出了控制单次审计分析中误标记操作比例的程序,仅要求正确分析可交换,而非拟合模型必须正确,并量化了当模型不完美或分析因内容被选中进行审查时,保证会被削弱多少。最后,本文确定了任何此类审计可报告的极限:低于一定幅度的错误完全无法被归因,因为它们与正确分析中的普通变异无法区分。该极限随收集到的正确分析数量增加而下降的速度极慢,以至于在当前使用的表示规模下,百倍的增加仅能将其降低不到2%,因此表示的维度而非训练数据的量是关键约束。

英文摘要

Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operations that depart from what that model predicts; how reliable such audits are has not been studied. This paper supplies that analysis. The choice of score determines whether an error can be localized at all. If each operation is scored by how surprising it is given the operation immediately preceding it, then operations that merely inherit an earlier error are indistinguishable from correct ones, so one mistake produces one flag; scores computed against a longer reconstruction of the intended analysis instead spread a single mistake across many operations. We quantify how far they spread, and how to choose the comparison length when an error accumulates gradually rather than at once. We then give procedures that control the proportion of falsely flagged operations within a single audited analysis, requiring only that sound analyses be exchangeable rather than that the fitted model be correct, and we quantify how much the guarantees weaken when the model is imperfect or when the analysis was selected for review in a way that depends on its content. Finally we establish a limit on what any such audit can report: errors below a certain magnitude cannot be attributed at all, being indistinguishable from ordinary variation among sound analyses. This limit falls so slowly as more sound analyses are collected that at the representation sizes now in use a hundredfold increase reduces it by under two percent, so the dimension of the representation rather than the volume of training data is the binding constraint.

发表机构

  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑