arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

如何在LLM裁判上运行统计并信任结果:使用evalstats对小样本AI评估进行校准推断

How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats

Ian Arawjo

arXiv 2609.35815首次发表:更新:

发表机构

Université de Montréal(蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对LLM裁判分数统计中的假阳性问题,提出校准推断方法,开发evalstats工具包,通过PPI和自举调整实现小样本AI评估的可靠统计。

AI 中文摘要

学术界的研究人员越来越多地基于LLM裁判分数和小样本AI评估来提出显著性声明。然而,如果没有良好校准的置信区间(CI)、假设检验和裁判偏差校正,这些声明是不可靠的。我们通过多项贡献来解决这些问题。首先,我们发现对原始LLM裁判分数进行统计会导致虚高的假阳性率:反直觉的是,对于许多评分者间一致性指标,假阳性风险在“几乎完美”的人机LLM一致性处达到峰值。为了帮助研究人员理解如何负责任地对LLM裁判进行统计,我们提供了混合人机AI裁判设计统计分析指南和工具,并通过预测驱动推断(PPI)实现了九种假设检验,包括四种基于秩的检验(Wilcoxon符号秩检验、Mann-Whitney U检验及其omnibus变体)的首个已知PPI校正。为了在小规模人工标注校准集下保持PPI++的稳定性,我们引入了自举自适应功率调整,该方法将估计权重向从标注数据估计的目标值收缩,并考虑该权重自身的采样方差。其次,通过蒙特卡洛模拟,我们得出了在小样本AI评估(N<100)中应使用何种CI、p值和FWER校正方法的建议,并警告研究人员不要使用自举CI。我们将这些建议打包到evalstats中,这是一个开源Python包,可自动选择校准方法,并在三个场景中进行了演示,其中一个场景中,一个在“实质性一致性”下验证的真实LLM裁判会导致研究人员发表虚假发现。evalstats可在https://this URL公开获取。

英文摘要

Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that running statistics over raw LLM judge scores leads to inflated false positives: counterintuitively, for many inter-rater agreement metrics, false positive risk peaks at "almost perfect" human-LLM agreement. To help researchers understand how to run statistics over LLM judges responsibly, we present guidance and tooling for the statistical analysis of mixed human-AI judge designs, and implement nine hypothesis tests via prediction-powered inference (PPI), including the first known PPI corrections for four rank-based tests (Wilcoxon signed-rank, Mann-Whitney U, and omnibus variants). To keep PPI++ stable with small human-labeled calibration sets, we introduce bootstrap-adaptive power tuning, which shrinks the estimated weight toward a target estimated from the labeled data, and accounts for that weight's own sampling variance. Second, through Monte Carlo simulations, we derive recommendations for what CI, p-value, and FWER correction methods to use for small-sample AI evaluations (N<100), and warn researchers against bootstrap CIs. We package these recommendations into evalstats, an open-source Python package that selects calibrated methods automatically, and demonstrate it in three scenarios, including one where a real LLM judge validated at "substantial agreement" would have led a researcher to publish a spurious finding. evalstats is publicly available at https://github.com/ianarawjo/evalstats.

Comments39 pages, 20 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑