arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CNM-BERT:基于表意文字描述序列的汉字可插入式结构嵌入

CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences

Thomas Sing-wing Wu, Liqian Yan

arXiv 2608.05167首次发表:更新:

发表机构

Shanghai Starriver Bilingual School; LinkScape(上海星河湾双语学校; LinkScape)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 CNM-BERT,通过将表意文字描述序列的结构嵌入融合到 BERT 中,提升了模型对稀有字和未登录字的理解,在结构探测基准及多个下游任务上均取得性能提升。

AI 中文摘要

基于 token 的编码器(如 BERT)将汉字视为原子标识符,忽略了其递归正字结构,导致模型依赖上下文共现,降低了对稀有字和未登录词(OOV)的性能。本文提出组合网络模型(CNM),一种轻量级增强方法,将离散的组合结构注入 Transformer 编码器。CNM 将表意文字描述序列(IDS)解析为树,通过递归 Tree-MLP 对其编码,并将结构嵌入融合到 BERT 中,无需修改主干。在 Wu 等人(2025)的结构探测基准上评估,CNM-BERT 在长尾字和未登录字上的结构准确率比最强基线 ChineseBERT 高 +9.8,部首 F1 值高 +7.7。此外,CNM-BERT 在 CLUE、MRC 和 NER 任务上,无论是 base 还是 large 规模,均取得一致提升,表明显式结构注入既实现了稳健的未登录词理解,又带来了切实的下游价值。

英文摘要

Token-based encoders like BERT treat Chinese characters as atomic identifiers, ignoring their recursive orthographic structure. Consequently, models rely on contextual co-occurrence, degrading performance on rare and out-of-vocabulary (OOV) characters. We propose the Compositional Network Model (CNM), a lightweight augmentation that injects discrete compositional structure into Transformer encoders. CNM parses Ideographic Description Sequences (IDS) into trees, encodes them via a recursive Tree-MLP, and fuses the structural embeddings into BERT without modifying the backbone. Evaluated on the Wu et al. (2025) structural-probing benchmark, CNM-BERT outperforms the strongest baseline (ChineseBERT) on long-tail and OOV characters by +9.8 Structure accuracy and +7.7 Radical F1. Furthermore, CNM-BERT achieves consistent gains across CLUE, MRC, and NER tasks at both base and large scales, demonstrating that explicit structural injection delivers both robust OOV understanding and tangible downstream value.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑