发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AUDITA是自主多智能体系统的审计层,结合防篡改智能体命令记录与可验证因果归因引擎,可减少责任错误、处理多重决定等场景,将责任判定从争论转为证据计算。
AI 中文摘要
物理自动化正朝着由AI大脑指挥的 embodied 机器集群扩展,早期部署已使工厂和仓库的生产效率远超任何人工生产线,且应用正在加速。但当它们的联合决策造成伤害时,各方(机器供应商、算法提供商、工厂运营商、保险公司、监管机构)都会相互指责,且目前没有方法能在它们之间划分责任。现有方法读取无法验证来源的日志并指定单一罪魁祸首,这歪曲了由多重决定、先发制人或遗漏导致的结果。我们提出AUDITA,这是一个审计层,将每个智能体间命令的防篡改记录与经过验证的分级因果归因引擎相结合。我们证明其裁决无法被操纵:遵守规则的智能体永远不会被判定有罪,转移责任的企图会被捕获并分级,且我们确定了基于证据的审计员可验证内容的精确限制。在实时语言模型管道上,它将标准法官基线的责任错误减少了约三倍;在基于事故结构的基准上,它在单一罪魁祸首基线失败的地方恢复了责任,且在伪造下保持不变。AUDITA将“谁该负责”的问题从关于日志的争论转变为基于证据的计算。
英文摘要
Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at production rates beyond any human line, and their adoption is accelerating. But when their joint decisions cause harm, everyone involved has reason to blame everyone else, the machine vendor, the algorithm provider, the factory operator, the insurer, and the regulator, and no method can divide the responsibility between them. Existing methods read logs whose origin they cannot verify and name a single culprit, misrepresenting outcomes that are overdetermined, preempted, or caused by an omission. We present AUDITA, an audit layer pairing a tamper-evident record of every inter-agent command with a certified, graded causal-attribution engine. We prove its verdict cannot be gamed: a rule-following agent can never be made to look guilty, an attempt to shift blame is itself caught and graded, and we establish the exact limit of what an evidence-based auditor can certify. On live language-model pipelines it reduces the standard judge baseline's responsibility error roughly threefold; on a benchmark of accident-grounded structures it recovers responsibility where single-culprit baselines fail, and stays invariant under forgery. AUDITA turns the question of who is to blame from an argument about logs into a calculation over evidence.