arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26319cs.CLcs.LG

噪声响应何时具有普适性?语言模型中作为隐藏变量的分词

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

Yefan Tao, Gerald Friedland, Luyang Kong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现语言模型对噪声的响应一致性取决于噪声规模,核心是训练目标而非架构,追溯到分词机制,可用于预测和提升模型抗噪声能力。

中文摘要 AI 辅助

文本神经模型在输入受到拼写错误、OCR错误或单词缺失等噪声干扰时,性能通常会下降。我们研究了句子嵌入模型和仅解码器大语言模型(LLM)在不同神经模型上的下降速率,发现其一致性取决于噪声规模:在单词级噪声下,架构差异极大的模型会沿几乎相同的曲线下降;而在字符级噪声下,它们则会出现分化。我们进一步确定决定因素是训练目标而非架构:涵盖六种预训练范式的八个编码器初始状态分散,经过简短的对比训练方法后会收敛到共同曲线。我们将单词/字符级噪声的差异追溯到分词:单个字符编辑会迫使分词器重新分割周围单词,对 token 序列的干扰远大于删除整个单词。这一发现及其潜在机制提供了实用方法,无需任何带噪声的评估即可预测模型的抗噪声能力,还可通过噪声增强训练在选定噪声规模下提升鲁棒性。

英文摘要

The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and decoder-only LLMs, and find that how consistent it is depends on the scale of the noise: under word-level noise, models with very different architectures decline along nearly the same curve, while under character-level noise they separate. We further identify the determining factor to be the training objective, not the architecture: eight encoders spanning six pretraining paradigms are scattered initially, and collapse onto a common curve after a short contrastive training recipe. We trace the word/character split to tokenization: a single character edit forces the tokenizer to re-segment the surrounding word, disturbing the token sequence far more than dropping a whole word does. This finding and its underlying mechanism provide a practical means to predict a model's robustness to noise without any noisy evaluation, and to install robustness at a chosen noise scale through noise-augmented training.

发表机构

  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑