发表机构
Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), VNU-HCM; Mohamed bin Zayed University of Artificial Intelligence(胡志明市理工大学计算机科学与工程学院(胡志明市国立大学所属); 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对越南占语、高棉语和岱侬语在自然语言处理中缺乏语料库的问题,引入CKTN语料库,指出现有多语言编码器不足,采用脚本感知适应方法,得到分割减少、分类性能强的编码器,揭示词汇重叠检索评估局限。
AI 中文摘要
越南少数民族语言在自然语言处理领域几乎缺失,挑战不仅在于数据稀缺。占语、高棉语和岱侬语在文字、与越南语的接触及标准化方面差异大,标准多语言适应可能学习错误信号。我们引入CKTN,首个针对这些语言的语料库和基准(44367个文档,2400万个子词标记),涵盖持续预训练、类别分类和摘要文档检索。现有多语言编码器严重分割这些语言,常见适应指标会误导。我们用脚本感知适应方法——词汇扩充结合校准替换词预训练,防止判别器利用琐碎脚本不匹配。结果是一个分割显著减少且分类性能最强的编码器,揭示了词汇重叠检索作为评估信号的局限性。
英文摘要
Vietnam's ethnic minority languages are almost absent from the field of Natural Language Processing (NLP), and the challenge goes beyond data scarcity: Cham, Khmer, and Tay-Nung differ sharply in script, Vietnamese contact, and standardization, conditions under which standard multilingual adaptation can learn the wrong signals. We introduce CKTN, the first corpus and benchmark for these languages (44,367 documents, 24M subword tokens), spanning continued pretraining, category classification, and summary-document retrieval. We show that existing multilingual encoders severely fragment these languages, and that common adaptation metrics can mislead: models may lower language-modeling loss or excel at lexical-overlap retrieval while still failing at semantic generalization across documents. We address this with a script-aware adaptation recipe - vocabulary augmentation combined with calibrated replaced-token pretraining - that prevents the discriminator from exploiting trivial script mismatches. The result is an encoder with substantially less fragmentation and the strongest classification performance among evaluated models, exposing the limits of lexical-overlap retrieval as an evaluation signal.