arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当分词器失效:面向低资源语言零样本迁移的字节级分块方法

When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages

Sanjeev Kumar, Atsuki Yamaguchi, Nikolaos Aletras

arXiv 2608.27658首次发表:更新:

发表机构

IIT Bombay; University of Sheffield(印度理工学院孟买分校; 谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对低资源语言处理中字节级模型的粒度不匹配问题,提出适配的分层网络框架,通过初始化字节嵌入、分块对齐损失及词性监督,在六种语言的词性标注任务上实现最高13.3%的性能提升。

AI 中文摘要

子词分词通过将主导语言的频率模式施加到同文字变体上,阻碍了低资源语言处理。字节级模型通过处理原始UTF-8字符规避了该问题,但在非拉丁文字的词级任务中会产生粒度不匹配。分层字节级架构通过将字节分组为词对齐分块解决了该不匹配,但这类架构需要大量训练数据,且与冻结的子词语言模型搭配时会出现表示错位。本文提出一种适配的分层网络框架,无需大量训练即可弥合该模态差距:该方法直接从冻结基础模型的子词表示初始化字节嵌入,应用分块对齐损失将动态分组的字节分块投影到预计算的子词目标,并插入轻量词性(POS)监督以引导边界检测。在六种语言上的实验表明,该无分词器方法可提升词级形态任务的性能,在词性标注上实现了最高13.3%的提升。

英文摘要

Subword tokenization hinders low-resource language processing by imposing frequency patterns from dominant languages onto script-sharing variants. Byte-level models bypass this issue by processing raw UTF-8 characters, yet they create a granularity mismatch for word-level tasks in non-Latin scripts. Hierarchical byte-level architectures address this mismatch by grouping bytes into word-aligned chunks. However, these architectures require massive training data and suffer from representational misalignment when paired with frozen subword-based language models. In this paper, we propose an adapted hierarchical network framework that bridges this modality gap without extensive training. Our method initializes byte embeddings directly from the subword representations of a frozen base model. We apply a chunk alignment loss to project dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech (POS) supervision to guide boundary detection. Experiments across six languages demonstrate that our tokenizer-free approach improves performance for word-level morphological tasks, yielding up to a 13.3% improvement on POS tagging.

CommentsAccepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑