arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全局一致性:当每个智能体都正确而团队仍然出错——多智能体协作的局部到全局语义基础

Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

Xin Heng

arXiv 2610.02036首次发表:更新:

发表机构

Tote AI(Tote AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出全局一致性问题,指出局部正确的智能体可能因共享状态缺失而团队出错,通过观测混淆不可能定理界定边界,并给出局部到全局的语义框架,证明局部智能无法替代全局状态。

AI 中文摘要

AI智能体各自可以做出局部有效的决策,但联合起来却可能产生无效的结果。我们将此称为全局一致性问题:这是共享状态的失败,而不仅仅是模型智能的失败。我们的观测混淆不可能定理给出了精确的边界。一个策略能够保证有效行动,当且仅当所有产生相同观测的世界共享一个可采纳的行动。如果k个不可区分的世界需要两两不相交的行动,那么最佳随机最坏情况成功率为1/k;更多的推理、角色、消息或样本无法恢复缺失的区分。更强的模型可以在其上下文内更好地推理,但它无法看到上下文之外的内容。随后我们给出了局部到全局的运行时语义X=(H, C, G, F; D):拓扑H记录重叠的范围;范畴C管理改变状态的行动;群胚G保留可逆的转换;层F测试局部视图是否能粘合为一个世界;最小历史D仅保留改变合法未来的区分。模型提出建议;执行框架拥有共享状态并管理提交。九项研究测试了该失败及其边界。在一个受控的修订基准上,当决策事件可见时,相同的前沿模型得分为40/40;当它被隐藏时,测试分支得分为12–17/40,与随机水平(1/3)一致;恢复一个权威事实后得分回到40/40。在TeamBench上,普通团队在5/5次运行中超出共享预算,可见的实时计数留下4/5次违规,而提交强制执行留下0/5次违规。在tau2-bench电信领域,静默回滚后当前状态检查得分为0.07,而执行框架得分为1.00。在传统求解器已经拥有完整相关状态的情况下,它与执行框架持平,正如预测的那样。反直觉的结论是,局部智能无法替代缺失的全局状态。

英文摘要

AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint actions, the best randomized worst-case success is 1/k; more reasoning, roles, messages, or samples cannot recover the missing distinction. A stronger model can reason better within its context, but it cannot see beyond it. We then give local-to-global runtime semantics X = (H, C, G, F; D): topology H records overlapping scopes; category C governs state-changing actions; groupoid G retains reversible translations; sheaf F tests whether local views glue into one world; and minimal history D keeps only distinctions that alter legal futures. Models propose; the harness owns shared state and governs commit. Nine studies test both the failure and its boundary. On a controlled revision benchmark, the same frontier model scores 40/40 when the deciding event is visible; when it is hidden, tested arms score 12--17/40, consistent with chance (1/3); restoring one authoritative fact returns 40/40. On TeamBench, ordinary teams exceed a shared budget in 5/5 runs, a visible live count leaves 4/5 violations, and commit enforcement leaves 0/5. In tau2-bench Telecom, current-state checks score 0.07 after silent reverts, while the harness scores 1.00. Where a conventional solver already owns the complete relevant state, it ties the harness as predicted. The counterintuitive conclusion is that local intelligence cannot substitute for missing global state.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑