arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22512cs.AI

HANSARD:面向自主多智能体AI系统的取证准备、运行时见证与分级归因的参考架构

HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems

  • Harokopio University of Athens(雅典哈罗科皮奥大学)
  • Democritus University of Thrace(色雷斯德谟克利特大学)
  • University of Western Macedonia(西马其顿大学)
  • Athena Research Center(雅典娜研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Christos Sardianos, Iliana Pla, Vasilis Efthymiou, Iraklis Varlamis, Thomas Lagkas, Panagiotis Sarigiannidis, Georgios Th. Papadopoulos

AI总结:

针对自主多智能体AI系统的归因洗钱问题,提出HANSARD参考架构,通过多阶段设计实现取证准备、运行时见证与分级归因,可检测归因洗钱并明确责任。

AI中文摘要:

自主多智能体系统如今已应用于金融、软件供应链和安全运营领域,目前已有首起由AI主导的入侵活动被报道。然而,当此类系统造成损害时,尚无方法能可靠地确定发生了什么、其原因是什么,或谁应承担责任。这是因为溯源取证的抽象层级不当,形式因果关系依赖于因果模型,而智能体审计则信任自记录。目标失效模式因此是归因洗钱,即把一项行为分散到多个冗余智能体中,直到没有任何一个智能体是必要条件原因。更糟的是,记录由嫌疑人生成,这构成了本研究全程采用的假设,智能体因此可能预判调查,且部分日志基础设施本身可能合谋。本文提出HANSARD,一种将问责制视为生命周期属性的参考架构:其一,运行前密封的准备配置文件约束后续发现可能的主张;其二,在智能体无法触及的五个 choke point 处捕获数据,使得遗漏和篡改都可被检测;其三,与PROV-DM对齐的类型化因果图在系统运行时累积,三个指标对其进行实时读取,以在无需裁决的情况下管控监督;其四,事件后重放产生经修正的Halpern-Pearl定义下的或有效果,以及补偿集规模;最后,协同剩余量衡量由组合而非个体导致的损害,使洗钱行为可见。随后将原因、责任和问责制分开报告,每项均受证据层级限制,同时还提供了未来研究议程。

英文摘要:

Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations. Already, the first largely AI-orchestrated intrusion campaigns have been reported. Yet, when such a system causes harm, no method can robustly establish what happened, what caused it, or who is accountable. This is because provenance forensics works at the wrong abstraction, formal causality assumes the causal model, and agent auditing trusts self-recording. The target failure mode is, thus, attribution laundering, i.e., spreading an act across redundant agents until none is a but-for cause. Worse, the record is produced by the suspects, which comprises the assumption adopted throughout this work. Agents may therefore anticipate the investigation and the part of logging infrastructure may itself collude. In this paper, HANSARD is proposed, a reference architecture treating accountability as a life-cycle property. First, a readiness profile sealed before operation bounds what later findings may claim. Second, capturing at five choke points beyond the agents' reach makes omissions detectable, not only tampering. Third, a typed PROV-DM-aligned causal graph accrues as the system runs, and three indicators read it live to gate oversight without adjudicating. Fourth, post-incident replay yields contingent effects under the modified Halpern-Pearl definition, together with a compensation-set size. Finally, a synergy residual measures harm due to the combination rather than to individuals, making laundering visible. Cause, responsibility and accountability are then reported separately, each capped by an evidentiary tier, while a future research agenda is also provided.

↑