发表机构
Shanghai Starriver Bilingual School; LinkScape(上海星河湾双语学校; LinkScape)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 CNM-BERT,通过将表意文字描述序列的结构嵌入融合到 BERT 中,提升了模型对稀有字和未登录字的理解,在结构探测基准及多个下游任务上均取得性能提升。
AI 中文摘要
基于 token 的编码器(如 BERT)将汉字视为原子标识符,忽略了其递归正字结构,导致模型依赖上下文共现,降低了对稀有字和未登录词(OOV)的性能。本文提出组合网络模型(CNM),一种轻量级增强方法,将离散的组合结构注入 Transformer 编码器。CNM 将表意文字描述序列(IDS)解析为树,通过递归 Tree-MLP 对其编码,并将结构嵌入融合到 BERT 中,无需修改主干。在 Wu 等人(2025)的结构探测基准上评估,CNM-BERT 在长尾字和未登录字上的结构准确率比最强基线 ChineseBERT 高 +9.8,部首 F1 值高 +7.7。此外,CNM-BERT 在 CLUE、MRC 和 NER 任务上,无论是 base 还是 large 规模,均取得一致提升,表明显式结构注入既实现了稳健的未登录词理解,又带来了切实的下游价值。
英文摘要
Token-based encoders like BERT treat Chinese characters as atomic identifiers, ignoring their recursive orthographic structure. Consequently, models rely on contextual co-occurrence, degrading performance on rare and out-of-vocabulary (OOV) characters. We propose the Compositional Network Model (CNM), a lightweight augmentation that injects discrete compositional structure into Transformer encoders. CNM parses Ideographic Description Sequences (IDS) into trees, encodes them via a recursive Tree-MLP, and fuses the structural embeddings into BERT without modifying the backbone. Evaluated on the Wu et al. (2025) structural-probing benchmark, CNM-BERT outperforms the strongest baseline (ChineseBERT) on long-tail and OOV characters by +9.8 Structure accuracy and +7.7 Radical F1. Furthermore, CNM-BERT achieves consistent gains across CLUE, MRC, and NER tasks at both base and large scales, demonstrating that explicit structural injection delivers both robust OOV understanding and tangible downstream value.