arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18284cs.CL

匈牙利制造:对生成式语言模型性能的评论

Made in Hungary: Comments on the performance of generative language models

  • Institute for Language Technologies and Applied Linguistics ELTE Research Centre for Linguistics(语言技术与应用语言学研究所 ELTE语言学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Mátyás Osváth, Enikő Héja, Noémi Ligeti-Nagy

AI总结:

本文评论匈牙利生成式语言模型研究,指出评估协议不可靠、训练流程欠佳及未评估能力损失等问题,强调严格实验设计的重要性。

AI中文摘要:

近年来,匈牙利出现了三项开发生成式语言模型的倡议。其背后的动机是相同的:对于匈牙利语,不存在具有给定能力的模型,或者现有的以英语为中心的模型能力有限。然而,对相应研究的详细审查揭示了若干方法论上的局限性。首先,评估协议的可信度值得怀疑。与Csibi等人[2026]的发现相反,在推荐的推理设置下进行评估显示,Qwen3-4B的得分高于其匈牙利语适配版本Racka-4B。数据污染在Yang等人[2025d]和Szentmihályi等人[2025]的工作中显而易见,可能使所报告的结果产生偏差。其次,训练流程在语料库整理和数据混合方面未达到当前最佳实践,这可能导致大量计算资源浪费在低质量数据上。缺乏受控的消融实验使得无法可靠评估这些选择。第三,三篇论文均未评估遗忘或能力损失。在原始基准测试的子集上测试适配模型表明,三种情况下的性能均有所下降,尤其是Racka-4B。鉴于语言模型开发涉及巨大的计算和财务成本,这些观察结果强调了严格实验设计的重要性。

英文摘要:

In recent years, three initiatives have emerged to develop generative language models in Hungary. The motivation behind them is the same. For Hungarian, no model with the given capability existed, or existing English-centric models offered limited proficiency. A detailed examination of the corresponding studies, however, reveals several methodological limitations. First, the reliability of the evaluation protocols is questionable. Contrary to the findings of Csibi et al. [2026], evaluation under the recommended inference settings shows that Qwen3-4B achieves higher scores than Racka-4B, its Hungarian-adapted version. Data contamination is evident in the work of Yang et al. [2025d] and Szentmihályi et al. [2025], potentially biasing the reported results. Second, the training pipelines fall short of current best practices in corpus curation and data mixture, which risks wasting substantial compute on low-quality data. The lack of controlled ablations prevents reliable assessment of these choices. Third, none of the three papers assessed forgetting or capability loss. Testing the adapted models on a subset of the original benchmarks indicates performance decline in all three cases, especially Racka-4B. These observations emphasize the importance of rigorous experimental design in language model development, given the significant computational and financial costs involved.

补充信息

↑