arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30373cs.CL

超越共识:用于主观评估的多智能体大语言模型评判器中的向下偏差与角色不对称

Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation

Minsoo Song, Chanwoo Kim, Sugyeong Eo, Chanjun Park

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现多智能体辩论(MAD)用于主观评估时存在角色不对称导致的向下偏差,会降低与人类判断的对齐度,消除角色不对称可恢复性能,揭示了共识式MAD协议的结构性局限。

中文摘要 AI 辅助

多智能体辩论(MAD)已被广泛用于改进基于大语言模型(LLM)的评估,方法是提示多个智能体进行协商并达成共识。然而,对于基于主观评分规则的评分而言,智能体间的一致性并不能保证与人类判断对齐。在本文中,我们将单评判器基线与基于共识的MAD协议在主观评估任务上进行比较,并设计了三项消融实验以分离角色提示、多轮交互和显式分数共享的影响。对六种LLM的评估显示,在六个评判模型中,单评判器基线平均实现了最强的人类对齐,而MAD在两项任务上均表现出人类对齐的下降。我们的消融实验表明,这种性能下降主要源于不对称角色提示,而非交互本身。具体而言,分配严格评判者角色会引入系统性的向下偏差,而共识过程无法纠正这种偏差。核心发现是,这种偏差反映了超出平均水平的严格立场主导:共识得分远超出独立严格条件与宽松条件的算术中点,而非对二者取平均。消除角色不对称(对称MAD)在很大程度上恢复了基线性能,而掩盖同伴分数则平均扩大了智能体间的分歧并恶化了平均人类对齐。这些发现表明,多智能体共识可能以牺牲真实人类对齐为代价强制实现人为一致性,揭示了用于主观评分的共识式、角色专业化MAD协议的结构性局限。

英文摘要

Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. However, for subjective rubric-based scoring, inter-agent agreement does not guarantee alignment with human judgments. In this paper, we compare a single-judge baseline against a consensus-based MAD protocol on subjective evaluation tasks and design three ablations to isolate the impact of role prompting, multi-round interaction, and explicit score sharing. Evaluations across six LLMs show that the single-judge baseline achieves the strongest human alignment on average across six judge models, whereas MAD shows degradation in human alignment on both tasks. Our ablations demonstrate that this performance drop stems primarily from asymmetric role prompting rather than the interaction itself. Specifically, assigning a strict judge role introduces a systematic downward bias that the consensus process fails to correct. The central finding is that this bias reflects strict-stance dominance beyond averaging: the consensus score falls well beyond the arithmetic midpoint of the standalone strict and lenient conditions, rather than averaging them out. Removing role asymmetry (Symmetric MAD) largely recovers baseline performance, while masking peer scores widens inter-agent disagreement on average and worsens average human alignment. These findings demonstrate that multi-agent consensus can enforce artificial agreement at the expense of true human alignment, revealing a structural limitation in consensus-style, role-specialized MAD protocols for subjective scoring.

发表机构

  • Soongsil University(崇实大学)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑