AI 中文总结
该研究针对 groundedness 验证测试多智能体小组交换意见可提升判断质量的假设,在6个基准上评估同质三智能体小组,发现其准确率差异在+8.5至-4.4个百分点,未证实该前提。
AI 中文摘要
大型语言模型(LLM)裁判正日益被组织为多智能体小组,前提是交换批评意见可提升判断质量。我们针对 groundedness 验证测试该前提,其中裁判需判断某主张是否有提供的证据支持。我们在6个公开事实验证和幻觉检测基准上评估了一个同质三智能体小组。相较于固定的单智能体参考,该小组的系统级准确率差异在+8.5至-4.4个百分点之间:两个数据集显示可靠提升,一个显示可靠下降,三个在统计上无结论。由于参考和小组使用不同的模型变体,这些差异表征了完整系统,而非孤立的因果辩论效应。
英文摘要
Large language model (LLM) judges are increasingly organized as multi-agent panels under the assumption that exchanging critiques improves judgment quality. We test this assumption for \emph{groundedness verification}, where a judge must determine whether a claim is supported by the supplied evidence. We evaluate a homogeneous three-agent panel on six public fact-verification and hallucination-detection benchmarks. Relative to a fixed single-agent reference, the panel's system-level accuracy difference ranges from $+8.5$ to $-4.4$ percentage points: two datasets show reliable gains, one shows a reliable loss, and three are statistically inconclusive. Because the reference and panel use different model variants, these differences characterize the complete systems rather than isolate a causal debate effect.