arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越原子词元:将音节分解用于语言模型预训练

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

arXiv 2609.21362首次发表:更新:

发表机构

University of Information Technology, Vietnam National University, Ho Chi Minh city; Mohamed bin Zayed University of Artificial Intelligence(胡志明市越南国立大学信息技术大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对传统分词忽视音节结构的问题,提出Phonemic Tokenizer,将音节分解为声母、韵母和声调,并构建PhonemicBERT,以极小词表在中文和越南语上取得与现有预训练模型相当或更优的性能。

AI 中文摘要

传统的分词器将文本表示为字符或统计得出的子词,忽略了音节的内部音系结构,且通常需要较大的词表。我们引入了具有语言学动机的分词器 Phonemic Tokenizer,用于越南语和中文,它将每个音节转换为国际音标(IPA),并将其分解为三个音系成分:声母(onset)、韵母(rime)和声调(tone)。这三个成分共同占据一个上下文位置,从而保持音节级的序列长度,同时实现音系相关音节之间的表示共享。非音系和不支持的单元通过字符级回退来处理。这种确定性的设计不需要依赖语料的词表学习,仅为中文生成112个词条、为越南语生成256个词条的词表。内在评估表明,该分词器在两种语言中均实现了显著更高的Rényi效率,以精确为1的生育率(Fertility)表示标准越南语音节词典中的每个词条,并且通常生成比现有预训练分词器更短的越南语序列。我们进一步将该分词器实例化为 PhonemicBERT,它结合了分解后的成分嵌入,并使用三个预测头重建完整的掩码音节。在受控的中文预训练设置下,PhonemicBERT-Zh 在各种语言理解任务上与字符、子词和 SubChar 替代方案相比具有竞争力或更优表现。PhonemicBERT-Vi 也取得了与成熟的越南语和多语言预训练模型相当或更优的结果。这些结果确立了音素分解作为原子和统计分段文本表示的紧凑、高效且可解释的替代方案。

英文摘要

Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequence length while enabling representation sharing across phonologically related syllables. Non-phonological and unsupported units are handled through character-level fallback. This deterministic design requires no corpus-dependent vocabulary learning and yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese. Intrinsic evaluation shows that the tokenizer achieves substantially higher Rényi efficiency in both languages, represents every entry in a standard Vietnamese syllable dictionary with a Fertility of exactly one, and generally produces shorter Vietnamese sequences than existing pretrained tokenizers. We further instantiate the tokenizer in \textbf{PhonemicBERT}, which combines factorized component embeddings and reconstructs complete masked syllables using three prediction heads. Under a controlled Chinese pretraining setup, PhonemicBERT-Zh is competitive with or outperforms character, subword, and SubChar alternatives across diverse language-understanding tasks. PhonemicBERT-Vi also achieves competitive or superior results to established Vietnamese and multilingual pretrained models. These results establish phonemic factorization as a compact, efficient, and interpretable alternative to atomic and statistically segmented text representations.

Commentsunder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑