发表机构
Bar-Ilan University; University of Cambridge; University of Washington; Allen Institute for AI(巴伊兰大学; 剑桥大学; 华盛顿大学; 艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出首个专为意第绪语构建的开源8B参数语言模型MameLoshnLM,通过自研语料库与基准优化Llama 3.1 8B,其表现优于同规模开源基线,为意第绪语NLP及低资源语言模型开发提供了基础与模板。
AI 中文摘要
我们推出MameLoshnLM,这是首个专为意第绪语构建的开源80亿参数语言模型。尽管意第绪语拥有丰富的文本传统,但其数字资源有限、可靠评估资源匮乏的现状,制约了意第绪语语言建模领域的发展。现有多语语料库和基准通常无法有效替代意第绪语,其中包含大量噪声、机器翻译及分类错误的文本。为解决这些缺口,我们推出Oytser——高质量意第绪语预训练语料库,它结合了当代原生网络资源与文学材料;还推出Kashes——涵盖翻译、语言分析、信息提取及语言理解的多任务基准。利用这些资源,我们对Llama 3.1 8B进行持续预训练,得到MameLoshnLM。在基准的各项任务中,MameLoshnLM的表现优于同等规模的开源基线模型。我们的分析显示,这些提升不仅是量化层面的:相较于通用多语模型,MameLoshnLM能更好地捕捉定义该语言的词汇与形态模式,这也揭示了噪声网络级多语数据对低资源语言存在的更广泛失效模式。我们的成果既为意第绪语自然语言处理提供了基础,也为历史底蕴深厚但数字资源匮乏的语言的模型开发提供了实用模板。
英文摘要
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
CommentsAccepted at the Conference on Language Modeling (COLM) 2026
Journal refProceedings of the Third Conference on Language Modeling (COLM 2026), 2026