arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更多辩论,相同证据:同质多智能体 groundedness 的结构局限

More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness

Yuelyu Ji

arXiv 2608.00243首次发表:更新:

AI 中文总结

该研究针对 groundedness 验证测试多智能体小组交换意见可提升判断质量的假设,在6个基准上评估同质三智能体小组,发现其准确率差异在+8.5至-4.4个百分点,未证实该前提。

AI 中文摘要

大型语言模型(LLM)裁判正日益被组织为多智能体小组,前提是交换批评意见可提升判断质量。我们针对 groundedness 验证测试该前提,其中裁判需判断某主张是否有提供的证据支持。我们在6个公开事实验证和幻觉检测基准上评估了一个同质三智能体小组。相较于固定的单智能体参考,该小组的系统级准确率差异在+8.5至-4.4个百分点之间:两个数据集显示可靠提升,一个显示可靠下降,三个在统计上无结论。由于参考和小组使用不同的模型变体,这些差异表征了完整系统,而非孤立的因果辩论效应。

英文摘要

Large language model (LLM) judges are increasingly organized as multi-agent panels under the assumption that exchanging critiques improves judgment quality. We test this assumption for \emph{groundedness verification}, where a judge must determine whether a claim is supported by the supplied evidence. We evaluate a homogeneous three-agent panel on six public fact-verification and hallucination-detection benchmarks. Relative to a fixed single-agent reference, the panel's system-level accuracy difference ranges from $+8.5$ to $-4.4$ percentage points: two datasets show reliable gains, one shows a reliable loss, and three are statistically inconclusive. Because the reference and panel use different model variants, these differences characterize the complete systems rather than isolate a causal debate effect.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑