arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25768cs.CLcs.AI

MoganBert-TR:采用CLM到MLM课程从头训练的土耳其编码器基础模型

MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay

AI总结:

本文提出从头训练的土耳其编码器基础模型MoganBert-TR,采用CLM到MLM两阶段课程学习,在TrGLUE等任务上优于同类模型,其衍生嵌入模型MoganBert-Embed在MTEB(土耳其)学生模型中排名第一

AI中文摘要:

土耳其编码器模型已采用现代架构,但预训练目标固定为掩码语言建模(MLM)。本文介绍MoganBert-TR,这是一个拥有1.49亿参数的土耳其编码器基础模型,基于特定语言过滤的语料库从头训练,以及由其衍生的嵌入模型MoganBert-Embed。MoganBert-TR在2373亿个 token 上进行训练,采用两阶段CLM到MLM课程:先进行因果语言建模(CLM),剩余阶段进行掩码语言建模,转换发生在词义消歧(WSD)调度的稳定阶段内。在相等步数预算下的受控消融实验中,该设计在土耳其MS MARCO检索任务上的性能比纯MLM高出2.7至3.7倍;测得的机制是嵌入几何,纯MLM下单个方向吸收28.1%的方差,而课程学习下为11.9%。随后,长上下文扩展和学习率衰减在共享前缀后分为两个分支:在1024上下文长度下运行衰减的最后部分,在五个配对随机种子上使TrGLUE平均得分提高0.49±0.26分(p=0.013),并以约4.3%的额外成本击败模型融合(model-soup)替代方案0.75分。MoganBert-TR在TrGLUE上得分78.41,是对比的土耳其ModernBERT模型中最佳的,在TabiBench上得分77.73,在8个类别中的2个类别领先,其中代码检索的优势最大(比TabiBERT高出3.62分)。通过教师蒸馏和多信号对比微调生成的MoganBert-Embed,在MTEB(土耳其)整体平均得分中排名学生模型第一,为68.30,且以小51倍的主干网络达到其75.7亿参数教师模型得分的99.5%。附带的50048 token 分词器在两个独立测试集上的压缩率和生成能力优于所有对比的土耳其分词器。权重、分词器、嵌入模型和评估代码:this https URL

英文摘要:

Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from scratch on a language-specifically filtered corpus, together with an embedding model derived from it (MoganBert-Embed). MoganBert-TR is trained over 237.3B tokens with a two-stage CLM-to-MLM curriculum: causal language modelling first, masked language modelling for the remainder, with the transition made inside the stable phase of a WSD schedule. In a controlled ablation under an equal step budget, this design outperforms pure MLM by 2.7-3.7x on Turkish MS MARCO retrieval; the measured mechanism is embedding geometry, where a single direction absorbs 28.1% of the variance under pure MLM against 11.9% under the curriculum. Long-context extension and learning-rate decay are then split into two branches after a shared prefix: running the final portion of decay at 1024 context improves the TrGLUE average by 0.49 +/- 0.26 points across five paired seeds (p = 0.013) and beats a model-soup alternative by 0.75 points at ~4.3% additional cost. MoganBert-TR attains 78.41 on TrGLUE, the best among the Turkish ModernBERT models compared, and 77.73 on TabiBench, where it leads two of the eight categories with the largest margin on code retrieval (+3.62 points over TabiBERT). MoganBert-Embed, produced through teacher distillation and multi-signal contrastive fine-tuning, ranks first among student models on the MTEB(Turkish) overall average with 68.30 and reaches 99.5% of its 7.57B-parameter teacher's score with a 51x smaller backbone. The accompanying 50,048-token tokenizer outperforms all compared Turkish tokenizers on compression and fertility across two independent test sets. Weights, tokenizer, embedding model and evaluation code: https://huggingface.co/moganai

补充信息

↑