AI 中文总结
针对多智能体辩论中虚构共识问题,提出主动溯源门控(APG)作为辩论后验证层,通过自我修正和严格门控提升数据溯源保真度,并在人类研究中获多数用户偏好,实现从被动记录到主动阻断的转变。
AI 中文摘要
基于大语言模型的多智能体辩论(MAD)系统正日益被用作分布式流程中的复杂决策管道。然而,其最终综合阶段仍缺乏充分控制。即使有详细的辩论日志,总结模型也容易生成流畅但虚构的辩论共识,这些共识并未基于辩论历史。为弥补这一安全缺口,本文开展实证研究,探讨引入主动的辩论后验证是否能减少此类缺乏事实支持的摘要的产生,同时仍提供有价值的信息。此外,还考察了在缺乏可靠妥协方案时,明确标示分歧是否更为可取。本文提出主动溯源门控(APG)作为辩论后验证层,将来源视为硬约束,分析辩论日志、审计每项声明并应用自我修正。在危机模拟中,自我修复机制在困难条件下将平均数据溯源保真度提升一倍以上,随后严格门控阻止无支持的声明并生成分歧报告。在人类研究中,绝大多数用户(超过75%)更倾向于在关键场景中明确报告失败,尽管他们中的大多数人认为基线系统生成的虚构共识更为流畅。我们的主要贡献是将数据来源追踪从被动记录转变为发布前的主动条件阻断。
英文摘要
Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are prone to fabricating smoothly written debate consensus that is not grounded in the debate's history. To address this safety gap, this paper presents empirical research and studies if the introduction of active post-debate verification can mitigate the production of such factually unsupported summaries, while still providing valuable information. Furthermore, it is examined whether explicitly signalling divergence is preferable in the absence of a reliable compromise. The Active Provenance Gate (APG) is introduced as a post-debate verification layer that treats the source as a hard constraint, analysing the debate logs, auditing each claim, and applying self-correction. In crisis simulations, the self-healing mechanism more than doubles the average data Provenance Fidelity in difficult condition scenarios, before the strict gate blocks unsupported claims and generates divergence reports. In the human study, a vast majority of the users (over 75%) preferred a report explicitly stating failure in critical scenarios, despite most of them perceiving fabricated consensus from the baseline system as more fluent. Our main contribution is the transition of data origin tracing from passive logging to active conditional blocking before publication.
CommentsAccepted for publication at the 38th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026)