arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化并缓解大语言模型中的朝鲜语谚文初声、中声、终声(Jamo)层面的排版漏洞

Quantifying and Mitigating Korean Jamo-Level Typographical Vulnerabilities in Large Language Models

Seojin Lee, Hwanhee Lee

arXiv 2608.30229首次发表:更新:

发表机构

Chung-Ang University(中央大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究量化了LLMs对朝鲜语Jamo层面排版扰动的脆弱性,提出TACoT方法,可在低推理成本下恢复思维链的大部分准确率增益。

AI 中文摘要

朝鲜语引入了普通字符级编辑模型未捕捉到的额外排版扰动层级:由于谚音节块内部由称为Jamo的子字符单元组成,键盘层面的错误可发生在一个音节内部,要么产生有效但语义改变的字符,要么在表面暴露原始Jamo。这两种结果都会破坏子词分词,且现有语法错误修正管道无法可靠纠正,使大语言模型(LLMs)直接暴露于损坏的输入中。为量化此漏洞,我们对KMMLU基准应用五种Jamo层面的扰动类型,并评估四个语言模型,发现准确率随扰动强度单调下降,且参数规模扩展无法赋予对抗音节内噪声的鲁棒性。我们进一步表明,受排版错误损坏的输入会在内部表示中诱导出与普通答案错误不同的明显偏移,且基于这些表示训练的简单线性探测器能以高AUROC检测未见过的扰动类型。受此信号启发,我们提出感知排版错误的思维链(Typo-Aware Chain-of-Thought,TACoT),仅当探测器检测到可能的排版错误时,才将输入路由至思维链推理,以一小部分推理成本恢复了思维链准确率增益的大部分。

英文摘要

Korean introduces an additional typographical perturbation level not captured by ordinary character-level edit models: because syllable blocks are internally composed of sub-character units called jamo, keyboard-level errors can occur within a syllable, either producing a valid but semantically altered character or exposing raw jamo on the surface. Both outcomes disrupt sub-word tokenization and are not reliably corrected by existing grammatical error correction pipelines, leaving LLMs directly exposed to corrupted inputs. To quantify this vulnerability, we apply five jamo-level perturbation types to the KMMLU benchmark and evaluate four language models, finding that accuracy declines monotonically with perturbation intensity and that parameter scaling does not confer robustness against intra-syllabic noise. We further show that typo-corrupted inputs induce a distinct shift in internal representations that is not reducible to ordinary answer incorrectness, and that a simple linear probe trained on these representations detects unseen perturbation types with high AUROC. Motivated by this signal, we propose Typo-Aware Chain-of-Thought (TACoT), which routes inputs to chain-of-thought inference only when the probe detects a likely typo, recovering a substantial portion of the CoT accuracy gain at a fraction of the inference cost.

CommentsAccepted to EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑