arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

情绪很重要:句法敏感性如何破坏安全对齐

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt, Simone Paolo Ponzetto

arXiv 2608.05409首次发表:更新:

发表机构

University of Mannheim; Technical University Clausthal(曼海姆大学; 克劳斯塔尔工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现大型语言模型存在句法敏感性漏洞,16个700亿参数级模型可通过调控句法特征触发或抑制拒绝行为,源于开源模型后训练数据的语言偏见,增加句法多样性可缓解问题。

AI 中文摘要

大型语言模型通常会经过后训练以使其与安全策略对齐,但存在许多复杂的越狱手段可规避已建立的安全措施。例如,Andriushchenko等人(2025)的前期研究发现,将语法时态从现在时改为过去时就足以引发有害响应。在本研究中,我们揭示了非祈使句法形式更普遍的失效问题。我们通过行为评估证明,这种句法漏洞存在于16个参数规模达700亿的模型中。为探究根本原因,我们应用因果中介分析,发现拒绝(即弃权不执行)部分取决于上游句法特征。通过调控这些纯句法特征,我们能够触发和抑制拒绝行为。最后,我们将这种不良条件归因于开源模型中存在语言偏见的后训练数据,并表明增加句法多样性可缓解该问题。我们的研究结果表明,当前的对齐方法引入了混杂变量,阻碍了拒绝决策的纯语义基础。

英文摘要

Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑