发表机构
Knight Foundation School of Computing and Information Sciences, Florida International University; SRI International; University of Florida; Oak Ridge National Laboratory; Army Cyber Institute, United States Military Academy; Oakland University(佛罗里达国际大学奈特基金会计算与信息科学学院; SRI国际公司; 佛罗里达大学; 橡树岭国家实验室; 美国军事学院陆军网络研究所; 奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Consilience框架,通过轮次保形校准程序保证多智能体通信控制的可靠性,在HiddenBench任务上提升了决策准确性与通信效率,优于现有协议甚至全信息基线。
AI 中文摘要
多智能体大语言模型(LLM)系统可通过汇集不同视角提升推理能力,但其有效性依赖于通信协调,尤其在隐式设定中——每个智能体仅持有正确决策所需的部分证据。现有协议(包括固定调度、轮次交换、非结构化辩论)无法保证对话行动的适当性。本文提出Consilience,一种推理时编排框架,可在分布式私有信息下引导并验证多智能体通信。每一轮中,Consilience会用紧凑状态(涵盖不确定性、分歧、证据增益、冗余及过早共识)总结讨论,随后选择通信干预方式(质疑、澄清、寻求证据或路由)及合适的发言者。其核心贡献是一种轮次保形校准程序,提供无分布、有限样本保证:每轮讨论中,在到达该轮的条件下,控制器提议行动的单步遗憾被校准阈值以至少1-α的边际概率约束;接受机制通过替换不可接受的提议,为执行的行动强制执行相同保证。在涵盖12种开放和闭源权重LLM的HiddenBench式隐式任务上,Consilience相比固定和非结构化讨论协议提升了决策准确性与通信效率,有时甚至超越全信息基线(所有智能体观察到全部证据)。这些结果表明,经过验证的自适应通信控制比增加信息可用性更具价值,为可靠多智能体LLM协调提供了实用机制。
英文摘要
Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness depends on coordinating communication, particularly in hidden-profile settings where each agent holds only part of the evidence required for a correct decision. Existing protocols, including fixed schedules, round-robin exchange, and unstructured debate, provide no guarantee that a conversational action is appropriate. We propose Consilience, an inference-time orchestration framework that both steers and certifies multi-agent communication under distributed private information. At each turn, Consilience summarizes the discussion using a compact state capturing uncertainty, disagreement, evidence gain, redundancy, and premature consensus, then selects both a communication intervention (challenge, clarify, seek evidence, or route) and an appropriate speaker. Its central contribution is a round-wise conformal calibration procedure that provides a distribution-free, finite-sample guarantee: at each discussion round, conditional on reaching that round, the one-step regret of a controller's proposed action is bounded by a calibrated threshold with marginal probability at least 1 - alpha; an acceptance mechanism enforces the same guarantee for the executed action by replacing inadmissible proposals. On HiddenBench-style hidden-profile tasks spanning 12 open and closed weight language models, Consilience improves decision accuracy and communication efficiency over fixed and unstructured discussion protocols, sometimes surpassing a full-information baseline where every agent observes all evidence. These results demonstrate that certified adaptive communication control can be more valuable than increasing information availability, providing a practical mechanism for reliable multi-agent LLM coordination.
Comments39 pages, 2 figures, 9 tables. Includes appendix with full proof of Proposition 1, expanded results, ablations, and complete prompt templates