发表机构
Tsinghua University; University College London; Autonavi(清华大学; 伦敦大学学院; 高德地图)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过谱残差多样性和分布误差两个指标,审计语言模型法官小组的有效规模,发现其因目标而异,不可互换作为小组质量或人类替代率的度量。
AI 中文摘要
一个语言模型小组代表了多少人类判断?答案取决于匹配的对象。我们对照经验性的人类标签分布来审计分类法官小组,保留了相对于单一金标签的二值错误所掩盖的异议。我们通过将归一化残差格拉姆矩阵的参与比与条件独立的人类参考抽取进行匹配来测量谱残差多样性,得到nu_H。我们分别匹配分布平方误差,得到nu_MSE。在三个ChaosNLI任务中,相同的32位法官小组具有nu_H=4.24--6.50,但nu_MSE=2.30--3.75。一个谱恒等式区分了决定误差的特征值、成员能量和平均方向权重。可实现硬标签小组表明,即使成员能量相等且相关性非负,更大的谱多样性也可能伴随更差的分布恢复。在观察到的小组中,组内大小排序一致性因任务而异;某些成员的添加产生冲突的变化,这些变化在两个项目半区中持续存在。共识方向在中心化残差方差中的份额为MNLI-m上的gamma_co=43.8%和SNLI上的33.7%,量化了平均保留的共享变异。我们提供了对齐的投票和分析协议,以审计这些区别。因此,有效大小是一种目标特定的测量:谱多样性和分布恢复不应被视为可互换的小组质量度量或一般的人类替代率。
英文摘要
A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives $ν_{\mathrm{MSE}}=2.304$, $3.750$, and $3.445$, whereas spectral matching gives $ν_H=4.242$, $6.459$, and $6.499$, a gap of $1.72$--$1.89\times$; a binary-error diagnostic credits the same panels with only $1.971$--$2.227$ effective votes. Extrapolating the distributional-error curve at fixed squared mean residual, mean member variance, and normalized mean covariance gives asymptotes of $2.392$, $3.990$, and $3.655$, with 32 judges already reaching $94.0$--$96.3\%$. An exact spectral identity explains the gap: error depends on member energy and on the orientation of residual variation relative to averaging, information that the participation ratio (PR) discards. A realizable hard-label construction confirms that higher spectral diversity can coexist with worse distribution recovery even under equal member energies and nonnegative correlations, and the consensus direction retains $γ_{\mathrm{co}}=43.8\%$, $33.7\%$, and $35.9\%$ of centered residual variance. An external check on CC-1000, a 1,000-item Civil Comments subset with a different panel, gives $ν_H=2.84$. For panel choice, we establish an existence result and one feasible path: exhaustive enumeration at $k\in\{5,7\}$ shows that panels beating the accuracy-top-$k$ baseline on both accuracy and $ν_H$ always exist, and greedily swapping at most two members reaches $24.8$--$56.0\%$ higher $ν_H$ at $0.10$--$1.10$ percentage points higher accuracy. Our dataset and code are available at https://github.com/Chao1208/32judges-votes.
Comments23 pages, 12 figures, and 13 tables. Code and data: https://github.com/Chao1208/chaosnli-judge-votes