arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

未经核验的审计:当LLM问责层转发而非检查时

Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check

Paul-Peter Arslan

arXiv 2609.07680首次发表:更新:

发表机构

Institute For Future Technologies(未来技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现多智能体LLM流水线的问责层在转发而非核验报告时,会因结论字段的锚定效应严重损害审计准确率,需确保证据独立性。

AI 中文摘要

多智能体LLM流水线日益跨越组织边界;当故障出现时,必须有人确定其来源。可获得的工件很少是完整的执行轨迹:它是每个智能体提交的报告,而提交的报告可以在其观察结果旁边陈述结论。使用一个预注册的、按机构划分的六智能体流水线,该流水线具有流程级信息边界、平衡的缺陷注入和匹配的干净孪生(每个链模型345,600个请求,两个模型),我们首先报告我们的预注册假设——即集体责任框架会随链长度降低升级——未得到支持。然而,该层仍然不对称地失败。它几乎不产生任何内容:在7,996个所有智能体都保持沉默的干净情节中,零指控。它过滤上游错误的能力很差,在智能体发出虚假警报的干净情节中,分别有34.4%和62.6%的情节点名无辜方。在没有任何智能体提出真正来源的条件下(一个链模型上59.5%的情节),阅读报告的审计员在4.1%的情况下恢复它——低于均匀猜测(20%)和最佳固定链接指控者(31.0%)——而从相同情节的原始文档中则达到60.3%。删除一个子句,即携带智能体自身结论的字段,在观察结果不变的情况下隔离了原因:准确率上升到45.2%(+41.2个百分点,95%置信区间+35.3至+46.9),而遵循率从94.4%崩溃到3.4%;在建议正确的情况下,同样的删除反而损失准确率,从70.5%降至55.7%。这种危害在四个条件中的四个条件下(+8.5至+39.0个百分点)在两个前沿审计员身上复制,并在第二个领域(+47.7和+61.1个百分点)中复制,而成本消失。净效应由上游可靠性以及两个条件幅度共同决定。问责层需要与其验证的结论充分独立的证据。

英文摘要

Multi-agent LLM pipelines increasingly span organisational boundaries; when a fault surfaces, someone must determine where it entered. The artifact available is rarely a full execution trace: it is the reports each agent filed, and a filed report can state a conclusion alongside its observations. Using a pre-registered, institutionally partitioned pipeline of six agents with process-level information boundaries, balanced defect injection and matched clean twins (345,600 requests per chain model, two models), we first report that our pre-registered hypothesis -- that collective responsibility framing degrades escalation with chain length -- is not supported. The layer nevertheless fails asymmetrically. It originates almost nothing: zero allegations across 7,996 clean episodes where every agent stayed silent. It filters upstream error poorly, naming an innocent party in 34.4% and 62.6% of clean episodes where an agent raised a false alarm. Conditional on no agent proposing the true origin (59.5% of episodes on one chain model), an auditor reading the reports recovers it in 4.1% of cases -- below a uniform guess (20%) and the best fixed-link accuser (31.0%) -- while reaching 60.3% from the raw documentation of the same episodes. Deleting one clause, the field carrying the agents' own conclusion, isolates the cause at constant observations: accuracy rises to 45.2% (+41.2 pp, 95% CI +35.3 to +46.9) and adherence collapses from 94.4% to 3.4%; where the suggestion was correct the same deletion instead costs accuracy, 70.5% to 55.7%. The harm replicates on two frontier auditors in four conditions out of four (+8.5 to +39.0 pp) and in a second domain (+47.7 and +61.1 pp), where the cost disappears. The net effect is governed by upstream reliability together with both conditional magnitudes. An accountability layer needs evidence sufficiently independent of the conclusions it verifies.

Comments18 pages, 5 figures. Pre-registration: https://doi.org/10.17605/OSF.IO/GBR3V. Code: https://github.com/Polpii/FaultLine

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑