arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

sk-bench:评估斯洛伐克语大语言模型的原生优先基准

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

Marek Šuppa, Ivan Vykopal, Andrej Ridzik, Kristián Sopkovič, Natália Kňažeková, Jaroslav Kopčan, Miroslav Blšták, Viktória Ondrejová, Daniel Hládek, Michal Gregor, Martin Tamajka, Marián Šimko

arXiv 2610.09152首次发表:更新:

发表机构

Comenius University in Bratislava; Kempelen Institute of Intelligent Technologies; Technical University of Košice; Cisco Systems(布拉迪斯拉发夸美纽斯大学; 肯佩伦智能技术研究所; 科希策技术大学; 思科系统公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对斯洛伐克语缺乏评估基准的问题,构建原生优先的sk-bench基准,含30个数据集,评估55个模型,发现原生数据、指令修复、测试时推理等关键经验,并开源数据与代码。

AI 中文摘要

多语言大语言模型基准测试遗漏了斯洛伐克语——一种拥有五百万使用者的形态丰富的西斯拉夫语,或仅通过机器翻译覆盖。我们提出sk-bench,一个原生优先的斯洛伐克语基准,包含十个技能类别下的30个数据集(33个评分任务变体)。其中11项资源是首次引入或打包用于生成式大语言模型评估,包括带有斯洛伐克语适配指令检查器的IFEval-SK,以及用于斯洛伐克语语法和形态学的原生Chiby/SKJ1资源。我们在同一测试框架下评估了55个开放权重和封闭权重模型。最佳开放模型落后专有API 12.6个百分点。对于原生和翻译的封闭式数据,模型排名相似(ρ≥0.98),尽管翻译对最强模型的区分度较低。相比之下,人工编写和LLM生成的问答问题对模型的排名不同(ρ=0.72)。对于Qwen3-14B,继续斯洛伐克语预训练使总体得分降低13.9个百分点。一个小的指令集恢复了该损失的四分之三。测试时推理将9B及以上模型的得分提高了8.5至12.5个百分点。这些发现共同为其他资源匮乏的语言提出了四个设计经验:在翻译失败处使用原生数据,在语言适应后规划指令修复,在扩展规模前启用测试时推理,并避免过度投资于目标语言提示。我们在以下网址发布数据和代码:此HTTPS URL。

英文摘要

Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($ρ\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($ρ=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench

CommentsAccepted to EMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑