发表机构
Guangdong Haiqixing Marine Technology Co., Ltd.(广东海启星海洋科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在微型中文BERT上受控比较MLM、WWM与MacBERT策略,发现MLM整体最优,WWM在困惑度和命中率上领先,而MacBERT因词典受限表现不佳,并揭示训练损失在混合替换下不可靠。
AI 中文摘要
预训练策略显著影响语言模型的质量,然而现有的掩码语言建模(MLM)、整词掩码(WWM)和MacBERT式替换的比较主要集中于基础规模模型(参数≥1.1亿)。本文在微型中文BERT模型(4层,隐藏维度256,参数870万)上对这三种策略进行了受控比较。在相同的架构、语料库(来自中文维基百科的129万句子)和超参数下,我们从零训练三个模型,并从五个内在维度进行评估:困惑度、MLM命中率、语义区分能力、语法判断能力和上下文敏感性。在微型规模下,MLM在整体内在性能上表现最佳(在5个维度中赢得3个),而WWM在困惑度(1.27对比2.10,提升39.5%)和MLM命中率(22%对比16%)方面均表现优异。值得注意的是,在严重受限的同义词词典(222个条目,覆盖率3.3%)下,MacBERT表现出严重的困惑度退化(47.23,比MLM高22倍),得出的排名(MLM > WWM >> MacBERT)与已确立的基础规模结论(MacBERT > WWM > MLM)显著不同。我们进一步识别出一个关键的评估陷阱:MacBERT实现了最低的训练损失(2.17),却拥有最高的困惑度(47.23),这揭示出在混合替换策略下,仅凭训练损失是不可靠的。所有模型和语料库均可在此https URL公开获取。
英文摘要
Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (>=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intrinsic performance (winning 3 of 5 dimensions), while WWM excels in both perplexity (1.27 vs. 2.10, a 39.5% improvement) and MLM hit rate (22% vs. 16%). Notably, MacBERT under a severely limited synonym dictionary (222 entries, 3.3% coverage) exhibits severe perplexity degradation (47.23, 22x higher than MLM), yielding a ranking (MLM > WWM >> MacBERT) that differs markedly from the established base-scale conclusion (MacBERT > WWM > MLM). We further identify a critical evaluation pitfall: MacBERT achieves the lowest training loss (2.17) yet the highest perplexity (47.23), revealing that training loss alone is unreliable under mixed replacement strategies. All models and corpus are publicly available at https://huggingface.co/eebyp.