arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31857cs.AIstat.AP

稀疏重叠下的LLM裁判验证:从推断到设计

LLM Judge Validation Under Sparse Overlap: From Inference to Design

Junxuan Li, Arko Mukherjee, Soumyabrata Pal

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明LLM裁判验证中重叠稀疏性是错误决策的主因,提出最小重叠公式和零成本分层分配方案,在四个评估矩阵上验证了10个裁判的有效性。

中文摘要 AI 辅助

验证LLM作为裁判需要估计其与人类的一致性,然而标注预算很少允许每个项目被多重标注。我们证明这种重叠稀疏性是错误部署决策的一阶决定因素:在5%的成对重叠率下,错误决策率达到25%,在十个候选者中选错最佳裁判的概率为65%。两个可行的杠杆是重叠的数量和分配。对于数量,我们推导出一个最小重叠公式,表明ρ≥0.25足以满足非边缘裁判,而边缘案例本质上仍然困难。对于分配,当分层信息丰富时,一种零成本的分层方案相对于随机抽样将错误拒绝率减半。我们在涵盖视觉评估、因果推理和摘要生成的四个评估矩阵上对10个LLM裁判进行了验证。

英文摘要

Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $ρ\geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.

补充信息

↑