arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32965cs.AIcs.SE

Relic:从多智能体协作到持久组织能力

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

Hongyi Du, Tianyi Zhang, Weijia Zhang, Yi Yang, Haofei Yu, Kunlun Zhu, Tianxiang Dai, Shang Jiang, Zhelun Gao, Jiaxin Pei, Shang Zhu, Jiaxuan You

首次发表
浏览论文内容

中文总结 AI 辅助

Relic将多智能体协作中的反复失败转化为组织可执行的协议,通过绑定触发器、责任和证据,提升完整合同交付率5.71个百分点,并超越同级基线,实现持久组织能力。

中文摘要 AI 辅助

在组织中,多个智能体常常会发生冲突:例如,一个编码智能体修改了仓库中的接口,而另一个智能体继续在旧版本上进行开发,导致现有测试变得过时。一次对话可以解决这一事件,但当参与者发生变化时,是什么让经验教训继续约束团队?我们引入了Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议。成员们反思可见的工作,提出规则,并管理其采纳过程。已采纳的协议将触发器、责任、所需证据和执行后果绑定到运行时,同时保持对修订和废止的开放性。在一个追踪案例中,反复的集成摩擦产生了一条接口审查规则,该规则约束了后续的拉取请求,并随着工作的继续而得到修订。在跨越十个软件工作负载和三个模型的360次受控运行中,与没有协议生命周期的匹配结构化团队相比,Relic将完整合同交付率从14.06%提高到19.76%(提升5.71个百分点),并在每个模型层级中改善了所有四个已验证的生产端点。在新成员迁移下,行为正确性在无继承协议时为25.4%,在提供相同规则作为可读文本时为34.6%,而在使用可执行绑定时为41.2%,比纯文本高出6.5个百分点。在完整的CooperBench基准上,排除损坏的基准对后,Relic达到367/477(76.9%),在同级结构化系统中建立了最佳报告结果。在固定的48对同模型子集上,Relic也超过了Solo(29/48对26/48),逆转了官方同级基线所展示的协调损失。这些结果共同表明,协作经验可以成为持久的组织状态,并且在其创造者之外仍然有用。

英文摘要

Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding 183 broken benchmark pairs, Relic achieves 371/469 (79.1%), establishing the best reported result among peer-structured systems. On the 47-pair same-model subset, Relic also exceeds Solo (28/47 vs. 26/47), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Harvey Mudd College(哈维穆德学院)
  • Yale University(耶鲁大学)
  • Stanford University(斯坦福大学)
  • Peking University(北京大学)
  • National University of Singapore(新加坡国立大学)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Together AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑