arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言环境和低资源语言中LLM-as-a-Judge的挑战与建议

Challenges and Recommendations for LLM-as-a-Judge in Multilingual Settings and for Low-Resource Languages

A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani

arXiv 2607.02235首次发表:更新:

发表机构

LT3, IDLab, Universiteit Gent, Barcelona Supercomputing Center, LMU Munich & Munich Center for Machine Learning, German Center for Addiction Research in Childhood and Adolescence, University Medical Center Hamburg-Eppendorf, Mila - Quebec AI Institute, McGill University, Canada CIFAR AI Chair(LT3、IDLab、根特大学、巴塞罗那超级计算中心、慕尼黑莱茵河大学及慕尼黑机器学习中心、德国成年期成瘾研究中心、汉堡埃彭多夫大学医学中心、魁北克人工智能研究所、麦吉尔大学、加拿大 CIFAR 人工智能主席)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文分析ACL论文中LLM-as-a-Judge在多语言和低资源语言评估中的应用,发现仅有33篇相关论文,存在评估结果不一致、过度信任LLM判断及依赖单一模型等问题,并提出改进建议。

AI 中文摘要

LLM-as-a-Judge已成为许多自然语言生成任务的主导评估范式,因为它克服了传统指标的缺陷,并且与人类判断高度相关,尽管这主要是在英语中。现在,人们尝试将LLM-as-a-Judge扩展到包括低资源语言在内的多语言环境。然而,LLM在低资源语言中的能力有限,并且在这些环境中通常缺乏足够的人类验证。为了突出问题的范围和当前实践,我们探索了ACL文集中关注多语言环境和低资源语言的各种任务的论文中LLM-as-a-Judge评估器的使用情况。在提及LLM-as-a-Judge的650篇论文中,只有33篇关注低资源或多语言环境。我们对这些论文的深入分析表明,评估结果不一致,存在在多语言环境中过度信任LLM判断的倾向,并且每项研究普遍依赖单一评判模型。为了进一步帮助NLP社区,我们最后提出了关于如何在多语言和低资源环境中使用LLM-as-a-Judge的建议。

英文摘要

LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to shortcomings of conventional metrics and high correlations with human judgment. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with a checklist of recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings

CommentsTo appear at EMNLP Findings 2026 (camera-ready version)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑