AI 中文总结
介绍 ELMOD 这个专为移动推理设计的 27 亿参数德语语言模型,在有限预算下用公开数据训练,开发特定德语数据预处理及质量过滤等步骤,使其在<3B 规模中性能最强,能与 7B 参数德语模型媲美。
AI 中文摘要
我们展示了 ELMOD(用于设备部署的高效语言模型),这是一个紧凑的(27 亿参数)德语语言模型,专为在资源受限的硬件上进行高效推理而设计。ELMOD 在有限的计算预算(55000 小时 H100 GPU)下仅使用公开可用数据进行训练。我们开发了一套特定于德语的数据预处理方法,其在处理形态变化、复合词和正字法惯例方面与面向英语的方法不同。此外,我们引入了质量过滤和重新措辞步骤,这提高了数据的教学质量,改善了退火阶段的性能,并降低了总体计算需求。由于我们的架构模型和数据选择,包括预过滤、提高教学质量的过滤和重新措辞,ELMOD 在其规模类别(<30 亿参数)中表现最强,与 70 亿参数德语模型的性能相当。
英文摘要
We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We developed a suite of German-specific data pre-processing, which differ from English-oriented counterparts in their handling of morphological variation, compounding, and orthographic conventions. Furthermore, we introduced a quality filtering and rephrasing step, which increased the instructional quality of the data, improved performance during the annealing phase, and reduced overall compute requirements. Thanks to our architectural model and data choices, including prefiltering, our educational-quality filtering and rephrasal to raise the educational-quality, ELMOD is the strongest performer in its size class (<3B), matching the performance of 7B-parameter models in German.
CommentsAccepted to KONVENS 2026