发表机构
Institute of IT, Kyrgyz State Technical University named after I. Razzakov; School of Computer Science, University of Windsor; St. Petersburg Department of the Steklov Math. Institute, RAS St. Petersburg State University(以I. 拉扎科夫命名的吉尔吉斯国立技术大学信息技术学院; 温莎大学计算机科学学院; 俄罗斯科学院斯捷克洛夫数学研究所圣彼得堡分部,圣彼得堡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对吉尔吉斯语大语言模型评估难的问题,利用KyrgyzLLM - Bench基准套件,包含两个本地数据集及多个翻译编辑后的数据集,在零样本和少样本设置下评估26个模型,分析性能等,发布相关资源以支持吉尔吉斯语NLP研究。
AI 中文摘要
评估跨语言的大语言模型具有挑战性,因为大多数多语言基准依赖翻译后的英语数据集,这往往会掩盖目标语言的语言和文化特异性。对于吉尔吉斯语等资源较少的语言,这个问题尤为突出,因为可靠的本地编写的评估数据稀缺。本文基于之前引入的吉尔吉斯语评估数据集,使用KyrgyzLLM - Bench基准套件首次对吉尔吉斯语的大语言模型进行了系统的大规模评估。KyrgyzLLM - Bench包括两个本地编写的数据集——KyrgyzMMLU和KyrgyzRC,以及WinoGrande、HellaSwag、BoolQ和TruthfulQA经过精心翻译和人工后编辑的版本。我们在零样本和少样本设置下评估了26个开源和闭源大语言模型,分析了模型性能、跨语言转移以及翻译工件对评估可靠性的影响。在不同的模型家族和任务中,模型排名在WinoGrande和BoolQ上从英语到吉尔吉斯语有广泛转移,在MMLU上转移程度较小,而HellaSwag表现出与翻译引起的合理性变化一致的显著英语 - 吉尔吉斯语性能差距。少样本提示在阅读理解方面提高了几个开源模型的性能,但在翻译任务中对专有模型的表现不一致。我们公开发布了所有数据集、评估代码和每个模型的结果,并将吉尔吉斯语任务集成到一个广泛使用的多语言评估框架中,以支持未来吉尔吉斯语自然语言处理的研究。
英文摘要
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two natively authored datasets$-$KyrgyzMMLU and KyrgyzRC$-$together with carefully translated and manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. We evaluate 26 open- and closed-source LLMs under zero-shot and few-shot settings, analyzing model performance, cross-lingual transfer, and the impact of translation artifacts on evaluation reliability. Across families and tasks, model rankings transfer broadly from English to Kyrgyz on WinoGrande and BoolQ, and to a lesser extent on MMLU, while HellaSwag exhibits a substantial English-Kyrgyz performance gap consistent with translation-induced plausibility shifts. Few-shot prompting improves several open-source models on reading comprehension but behaves inconsistently for proprietary models on translated tasks. We publicly release all datasets, evaluation code, and per-model results, and integrate the Kyrgyz tasks into a widely used multilingual evaluation framework to support future research on Kyrgyz NLP.
CommentsPreprint; manuscript currently under consideration at a journal