发表机构
AO, Inc.; OrbLabs AG(AO公司; OrbLabs AG)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体LLM谈判中与有状态对抗性守门人的问题,提出含5支柱宪法、4层群体和认知退火的控制栈,实验表明其可高效从死锁恢复,且不增加额外模型成本。
AI 中文摘要
与有状态对手谈判的多智能体大语言模型(LLM)系统在三个方面浪费模型调用:永远无法满足对手隐藏接受条件的礼貌循环、触发重试的格式错误输出,以及对手要求智能体必须拒绝的合规死锁。我们研究了一个由三部分组成的控制栈——5支柱运行时宪法、4层群体(主管、三智能体多数投票、监控器、模式硬门)和认知退火(确定性死锁检测、智能体侧上下文的原子清除、规范恢复消息)——以对抗一个已发布的对抗性守门人(Gatekeeper),其接受规则为固定正则表达式,且其LLM仅生成回复文本。测试平台有一个已知解决方案:它测量控制栈是否针对群体漂移执行符合宪法的策略并从死锁中恢复,而非它是否发现任何新内容。在每种配置下进行5次运行(共30次运行;智能体为Gemini 2.5 Pro,守门人为Claude Haiku 4.5),我们发现:(i)宪法和主管使可接受的框架成为可能但不可靠——基线为0/5解锁,加入宪法后为1/5和2/5;当群体解锁时,它在1轮内完成,调用次数为7-8次,约15k令牌(比基线低67-73%);当未解锁时,成本增加17-38%;(ii)监控器和硬门不减少解锁次数,但会留下审计跟踪;(iii)在蜜罐到合规的死锁下,仅LLM引导的5次运行中0次逃脱,而原子清除加规范恢复消息的5次运行全部逃脱(Fisher p=0.008),且调用预算相同,恢复消息的调用次数为0;LLM编写的恢复消息在确定性预检查中5次全部失败,尽管LLM监控器已批准其中4次。关于平均调用和令牌减少的预注册假设未得到支持。每个分支的成本都受确定性停止规则的限制;控制栈在不增加额外模型成本的情况下添加了恢复机制。
英文摘要
Multi-agent LLM systems negotiating with a stateful counterpart waste model calls in three ways: polite loops that never meet the counterpart's hidden acceptance condition, malformed outputs that trigger retries, and compliance deadlocks in which the counterpart demands something the agent must refuse. We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text. The testbed has a known solution: it measures whether the stack executes a constitution-aligned strategy against swarm drift and recovers from deadlock, not whether it discovers anything. In five runs per configuration (30 runs; Gemini 2.5 Pro agents, Claude Haiku 4.5 Gatekeeper) we find: (i) the constitution and Director make an acceptable framing possible but not reliable - 0/5 baseline unlocks versus 1/5 and 2/5 with the constitution; when the swarm unlocks it does so in one turn with 7-8 calls and about 15k tokens (67-73% below baseline); when it does not, it costs 17-38% more; (ii) the Monitor and hard gate do not reduce unlocks and leave an audit trail; (iii) under a honeytrap-to-compliance deadlock, LLM-only steering escapes 0 of 5 times while atomic purge plus a canonical strike escapes 5 of 5 (Fisher $p = 0.008$) at the same call budget, with zero calls for the strike. LLM-written strikes failed the deterministic pre-flight 5 of 5 times although an LLM Monitor had approved four. Pre-registered hypotheses on average call and token reduction were not supported. Cost is bounded in every arm by deterministic stop rules; the stack adds recovery at no extra model cost.
Comments29 pages, 3 figures. The Gatekeeper, agents, constitution, lexicon, 30 run logs and analysis scripts are released (see Appendix F). Companion paper: arXiv:2610.09772