arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02941cs.CL

形式对齐而非意义对齐:低资源孟加拉语贬损言论中大型语言模型(LLM)安全的理解-遏制解耦

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

Shadab Bin Habib, A K M Ferdous Reza Habib, Subarno Neel, Adib Sakhawat

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对5个前沿LLM,在6种协议下对孟加拉语贬损言论审计,验证了LLM安全的理解-遏制解耦假设,发现高资源基准无法保障低资源安全,需基于意义的遏制。

中文摘要 AI 辅助

我们针对5个前沿大型语言模型(LLM),在6种协议下对原生孟加拉语贬损言论(gali)进行审计,以验证单一假设:理解-遏制解耦。我们提出,当前的安全对齐绑定于高资源表面形式而非有害意义,导致模型理解低资源 slur( slur 指贬损性词语)的能力与遏制其的能力相互独立。所有协议均与人工校准基线(kappa=0.84)一致地证实了该假设。在基线水平下,模型在孟加拉语中表现出7.92个百分点的理解缺陷,同时在两种语言中保持相同的92.83% token( token 指模型处理的文本单元)泄漏率。严重程度校准追踪表面解剖线索而非组合性危害(温和俚语误差+4.00,威胁误差-2.00),而正字法扰动下的表观遏制增益实则是分词器驱动的“遏制幻象”。关键的是,显式思维链(Chain-of-Thought)推理可挽救理解(通过率94.72%),同时系统性瓦解遏制(使用率96.23%);此外,专家角色(expert-persona)框架将拒绝率降至6.57%,表明基于关键词的过滤器完全忽略了非人化的社群贬损语。我们的发现证明,高资源基准无法验证低资源安全,需要基于意义的遏制。

英文摘要

We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesis against a human-calibrated baseline (kappa = 0.84). At baseline, models exhibit a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. Severity calibration tracks surface anatomical cues over compositional harm (+4.00 error on mild slang; -2.00 on threats), while apparent containment gains under orthographic perturbation prove to be a tokenizer-driven "containment mirage." Crucially, explicit Chain-of-Thought reasoning rescues comprehension (94.72% Pass) while systematically dismantling containment (96.23% Use). Furthermore, expert-persona framing collapses refusal to 6.57%, revealing that keyword-based filters ignore dehumanizing communal slurs entirely. Our findings demonstrate that high-resource benchmarks cannot certify low-resource safety, necessitating meaning-grounded containment.

发表机构

  • Islamic University of Technology(伊斯兰科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑