AI 中文总结
针对NLP模型人工评估成本高、效率低的问题,将其形式化为多臂老虎机最佳臂识别问题,提出自适应采样算法,可优化预算分配,提升评估效率与模型区分度。
AI 中文摘要
尽管人工评估是许多自然语言处理(NLP)任务中的黄金标准,但它存在成本过高、可扩展性差的问题。在识别表现最佳的模型时,典型的评估协议会通过在整个基准测试中详尽评估所有模型来浪费精力,这是一种安全但低效的方法。在本研究中,我们将多模型人工评估形式化为相关臂的多臂老虎机设置中的最佳臂识别问题,其中拉动一个臂对应于对一个模型进行人工评估。通过基于迄今为止获得的中间模型排名自适应采样,我们可以将标注预算集中在最具竞争力的模型上。我们证明了所提出算法的最优性,并表明其提高了对表现最佳模型之间的区分度,这使得评估更快、更便宜,且更符合大规模竞赛评估目标。
英文摘要
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.
CommentsEMNLP 2026; typeset with Typst