发表机构
Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM评判器预测人类判断分布(软标签)性能差的问题,提出NAPHA方法,可提升软标签预测性能,尤其对高熵实例效果显著。
AI 中文摘要
大语言模型作为评判器(LLMaJ)框架为自动评估提供了一种高性价比且可复现的解决方案。然而,当前评估实践通常将LLMaJ的评判与聚合的真实标签进行比较,忽略了人类标签变异(HLV)中包含的宝贵信息。受越来越多利用HLV的研究工作启发,我们系统研究了LLMaJ在预测单个聚合的真实硬标签以及代表人类判断分布(HJD)的非聚合软标签方面的性能。我们在五个不同数据集上的结果表明,尽管大型语言模型在大多数任务的硬标签预测上达到了接近人类的性能,但它们在预测软标签时表现不佳。为解决这一局限,我们提出了NAPHA(熵感知事后对齐),这是一种简单而有效的轻量级事后对齐方法,通过首先将实例分配到离散熵类,然后将其路由到专门训练的对齐模型,使LLM分布与HJD匹配。我们发现,NAPHA在基础LLM模型和数据集上始终提升了软标签预测性能,在高熵实例上的提升尤为显著,而高熵实例正是捕捉不同人类观点最为关键的地方。我们还通过神谕实验表明,改进熵类预测可大幅提升NAPHA的实际有效性。
英文摘要
The LLM-as-a-judge (LLMaJ) framework offers a cost-effective and reproducible solution for automatic evaluation. However, current evaluation practices typically compare LLMaJ judgments against aggregated ground-truth labels, overlooking the valuable information contained in Human Label Variation (HLV). Inspired by an increasing line of work that proposes to leverage HLV, we systematically study LLMaJ performance on predicting both a single, aggregated ground truth hard-label and unaggregated soft-labels that represent Human Judgment Distributions (HJD). Our results across five diverse datasets reveal that while LLMs achieve near human-level performance at hard-label prediction on most tasks, they exhibit poor performance when predicting soft-labels. To address this limitation, we propose NAPHA (eNtropy-Aware Post-Hoc Alignment), a simple yet effective lightweight post-hoc alignment method that matches the LLM distribution to the HJD by first assigning an instance to a discrete entropy class and then routing it to specialized, trained alignment models. We find that NAPHA consistently improves soft-labels prediction across base LLM models and datasets, with particularly strong gains on high-entropy instances where capturing diverse human perspectives is most critical. We also show via oracle experiments that improving entropy class prediction can substantially enhance NAPHA's practical effectiveness.
CommentsAccepted to EMNLP 2026