arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03425cs.MAcs.AI

文明框架:以主权为锚的个人多智能体系统间通信

The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems

Guangjun Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出以人类为主权者的文明框架及大使馆协议解决多智能体通信的上下文丢失问题,通过1908次试验发现时间权重效应的隐患,相关结果为探索性,计划开展复制实验。

中文摘要 AI 辅助

人类是AI系统之间的传输层,每一次交互都会丢失上下文。我们提出了文明框架,其可寻址主体是文明而非智能体(包含一位人类主权者、一个持久账本和可互换的智能体),以及大使馆协议,一种与载体无关的覆盖网络:消息异步到达接收方的常驻账本端点,接收方的任何在线智能体均可处理,且以双方账本上的承诺状态而非传递作为基准事实。权威源于记忆:智能体代表其文明行事的权力受限于其可访问的记忆,并通过签名凭证外化,这与文明层面的声誉相分离。我们识别出时间权重效应,这是AI间通信中的一种隐患,即先到达的内容会获得不应有的权威,并在一个前沿模型中通过一项预注册的1908次试验实验对其进行测试。在移除验证的情况下,先到达的错误上游声明捕获了54.2%的答案(在完全验证下为4.2%),而当接收方已密封其自身答案后到达的相同声明捕获了31.6%(两个提示壳的长度不匹配,因此部分差距可能反映了壳的形式;见第7节),且两个预注册的问题集规范对这两个结论达成一致(排除规范被预注册为效力不足)。两个次要结果,即来自指令级来源标记的缓解效果和密封答案的准确性等价性,依赖于规范,仅在所有问题规范下成立。由于对工具使用的预注册检查未通过其调用预算条件,该注册将此轮归类为不确定,上述所有结果(主要和次要)均被报告为探索性;计划进行一项由工具强制预算的复制实验。该框架的文明内部层已具备可用实现。

英文摘要

Humans are the transport layer between AI systems, losing context at every hop. We present the Civilization Framework, whose addressable party is the civilization, not the agent (one human sovereign, a persistent ledger, and interchangeable agents), and the Embassy Protocol, a carrier-agnostic overlay: messages arrive asynchronously at a resident ledger endpoint, any online agent of the receiver handles them, and commitment state on both ledgers, not delivery, is ground truth. Authority derives from memory: an agent's power to act for its civilization is capped by the memory it can access and externalized through signed credentials, separate from civilization-level reputation. We identify the temporal-weight effect, a hazard in AI-to-AI communication where what arrives first acquires unearned authority, and test it in one frontier model in a preregistered 1,908-trial experiment. With verification removed, an incorrect upstream claim arriving first captures 54.2% of answers (4.2% under full verification), while the same claim arriving after the receiver has sealed its own answer captures 31.6% (the two prompt shells are not length-matched, so part of that gap may reflect shell form; see Section 7), and both registered question-set specifications agree on these two verdicts (the exclusion specification is preregistered as under-powered). Two secondary results, the mitigation from instruction-level provenance labeling and sealed-answer accuracy equivalence, are specification-dependent, holding only under the all-questions specification. Because a registered check of tool use failed its call-budget condition, the registration classifies the round as inconclusive and every result above, primary and secondary, is reported as exploratory; a replication with harness-enforced budgets is planned. The framework's intra-civilization layer has a working implementation.

发表机构

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑