发表机构
NVIDIA; COPA(英伟达; COPA)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对亚美尼亚语低资源问题,整理发布ArmWeb、ArmSTEM数据集,预训练得到首个带完整数据和方法的开源亚美尼亚语LLM arm-gemma-e4b,验证了STEM数据可弥补新闻预训练的知识损失,公开全部资源。
AI 中文摘要
亚美尼亚语是一种形态丰富的低资源语言,其预训练数据匮乏,且尚未有开源亚美尼亚语大语言模型(LLM)随可复现所需的数据和方法发布。为解决这一缺口,我们整理并发布了两个数据集:ArmWeb是经广泛验证的437万份亚美尼亚语文档语料库;ArmSTEM是包含37.3万道数学与科学题目的英-亚平行语料,题目及分步解答已被译为亚美尼亚语,并通过保留答案的LLM判断和人工评估双重验证。在这些数据集上对Gemma-4-E4B进行持续预训练得到arm-gemma-e4b,该模型性能优于所有现有开源亚美尼亚语模型及其未适配的基础模型,是首个拥有完整训练数据和方法的开源亚美尼亚语LLM。我们的消融实验表明:仅基于新闻数据的持续预训练可提升流畅度但会损害知识,这一模式在现有亚美尼亚语模型中也存在;而少量经验证的翻译后STEM数据可逆转该损失。我们还发现,最大的公开亚美尼亚语语料库与网络衍生的评估面板高度重叠,包括FineWeb-2内部存在训练-测试自重叠。我们公开发布所有数据、模型及代码。
英文摘要
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.
Comments18 pages, 4 figures, 13 tables. Data and model: https://huggingface.co/collections/COPA-AI/armenian-llm-ecosystem. Code: https://github.com/COPATeam/armenian_llm_ecosystem