发表机构
Comenius University in Bratislava; Kempelen Institute of Intelligent Technologies; Technical University of Košice; Cisco Systems(布拉迪斯拉发夸美纽斯大学; 肯佩伦智能技术研究所; 科希策技术大学; 思科系统公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对斯洛伐克语缺乏评估基准的问题,构建原生优先的sk-bench基准,含30个数据集,评估55个模型,发现原生数据、指令修复、测试时推理等关键经验,并开源数据与代码。
AI 中文摘要
多语言大语言模型基准测试遗漏了斯洛伐克语——一种拥有五百万使用者的形态丰富的西斯拉夫语,或仅通过机器翻译覆盖。我们提出sk-bench,一个原生优先的斯洛伐克语基准,包含十个技能类别下的30个数据集(33个评分任务变体)。其中11项资源是首次引入或打包用于生成式大语言模型评估,包括带有斯洛伐克语适配指令检查器的IFEval-SK,以及用于斯洛伐克语语法和形态学的原生Chiby/SKJ1资源。我们在同一测试框架下评估了55个开放权重和封闭权重模型。最佳开放模型落后专有API 12.6个百分点。对于原生和翻译的封闭式数据,模型排名相似(ρ≥0.98),尽管翻译对最强模型的区分度较低。相比之下,人工编写和LLM生成的问答问题对模型的排名不同(ρ=0.72)。对于Qwen3-14B,继续斯洛伐克语预训练使总体得分降低13.9个百分点。一个小的指令集恢复了该损失的四分之三。测试时推理将9B及以上模型的得分提高了8.5至12.5个百分点。这些发现共同为其他资源匮乏的语言提出了四个设计经验:在翻译失败处使用原生数据,在语言适应后规划指令修复,在扩展规模前启用测试时推理,并避免过度投资于目标语言提示。我们在以下网址发布数据和代码:此HTTPS URL。
英文摘要
Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($ρ\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($ρ=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench
CommentsAccepted to EMNLP 2026 Main