arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SoK:去中心化智能体经济基础设施

SoK: Decentralized Agent Economic Infrastructure

Rui Sun, Xihan Xiong, Qin Wang, Fei Gao, Zelin Li, Zehua Cheng, Jiahao Sun, Zhipeng Wang

arXiv 2610.01756首次发表:更新:

发表机构

Newcastle University; University of Bristol; CSIRO; University College London; Ohio State University; University of Oxford; The University of Manchester(纽卡斯尔大学; 布里斯托大学; 联邦科学与工业研究组织; 伦敦大学学院; 俄亥俄州立大学; 牛津大学; 曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文系统化研究了去中心化智能体经济中跨协议组合导致端到端保证失效的问题,提出保证闭合标准,并通过大规模实验揭示验证与结算间的失败模式及修复方向。

AI 中文摘要

去中心化智能体经济日益将单个任务构建在分别设计和保护的各种协议之上。这产生了一个简单的问题:工作流在每一步看起来都是正确的,但仍可能产生错误的结果。例如,一个正确的托管(escrow)可能会在授权批准时释放付款,而该批准几乎不能证明交付的工作确实满足了任务要求。我们在智能体任务的整个生命周期中系统化了这一问题。我们的研究将安全和经济需求组织为六个阶段中的17个属性族,并分别评估收据健全性(receipt soundness)和完整性(completeness)。我们考察了12个系统和标准、5个可复用机制族以及4个经典基线。我们引入了保证闭合(guarantee closure),这是一种任务相对标准,用于确定在一个阶段建立的保证是否仍然可用,并约束依赖这些保证的后续决策。我们将该标准应用于受控工作流和原生工作流,覆盖了840次匹配执行,并在一个有限目标-任务域上进行了详尽的11,648例检查。我们的结果揭示了验证与结算之间反复出现的失败,即符合要求的工作可能仍不被接受,或者有效的证据可能被忽略。公开记录和模型判断进一步区分了记录的批准与任务符合性的证据,而经济分析则识别了这些保证背后的报告、惩罚和共享错误假设。这些发现显示了端到端保证在何处失败,以及需要修复什么才能在工作流中保持这些保证。

英文摘要

Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides little evidence that the delivered work actually satisfied the task. We systematize this problem across the full lifecycle of an agent task. Our study organizes security and economic requirements into 17 property families over six stages, with receipt soundness and completeness assessed separately. We examine 12 systems and standards, five reusable mechanism families, and four classical baselines. We introduce guarantee closure, a task-relative criterion for determining whether guarantees established at one stage remain available and constrain the later decisions that depend on them. We apply the criterion to controlled and native workflows, covering 840 matched executions and an exhaustive 11,648-case check over a finite objective-task domain. Our results expose recurring failures between verification and settlement, where conforming work can remain unaccepted or valid evidence can be ignored. Public records and model judgments further distinguish recorded approval from evidence of task conformance, while economic analysis identifies the report, penalty, and shared-error assumptions behind these guarantees. These findings show where end-to-end guarantees fail and what must be repaired to preserve them across the workflow.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑