LLM对齐——语义保持变换下的效用不对称性
LLM Alignment--Utility Asymmetry under Semantic-Preserving Transformations
浏览论文内容
中文总结 AI 辅助
本研究通过合成语义保持变换,探针测试LLM对齐泛化,发现对齐-效用不对称性:模型在变换下效用保留而对齐失败率急剧上升。
中文摘要 AI 辅助
大语言模型(LLM)对齐旨在确保模型保持有用性和安全性,但其在输入分布偏移下的稳定性尚未被完全理解。先前的研究表明,对齐模型可能在越狱提示、替代编码和跨语言迁移下失效,但这些失败通常被作为攻击来研究,而非作为对齐泛化的受控探针。此外,现有证据主要基于预训练期间已表示的自然语言变异,这留下了未解决的问题:对齐是否随语义内容泛化,还是仍与表面模式紧密相关。在本文中,我们使用基于规则且可逆的合成语义保持变换来研究这一问题,这些变换在保留任务相关含义的同时,将输入偏移到标准语言变异之外。在四个开源和四个商业模型中,在微调和上下文学习两种设置下,我们使用这些变换作为对齐泛化的探针,并识别出一种经验模式,我们称之为“对齐-效用不对称性”:一旦模型能有效处理变换后的输入,任务效用通常大幅保留,而对齐失败则更急剧增加。例如,适应的GPT-4.1 mini在变换下仅表现出有限的效用下降,而其有害率从13.3升至74.3;Gemini 3 Flash同样保持接近原始的效用,而其有害率从2.3增至43.0。综合来看,这些结果表明,语义保持的分布偏移可能暴露当前LLM中效用与对齐泛化之间的一个反复出现的差距。
英文摘要
Large Language Model (LLM) alignment is intended to ensure that models remain helpful and safe, but its stability under input distributional shift is not yet fully understood. Prior work shows that aligned models can fail under jailbreak prompts, alternative encodings, and cross-lingual transfer, yet these failures are usually studied as attacks rather than controlled probes of alignment generalization. Moreover, existing evidence is largely grounded in natural language variation already represented during pretraining, leaving unresolved whether alignment generalizes with semantic content or remains tied to superficial surface patterns. In this paper, we study this question using synthetic semantic-preserving transformations that are rule-based and invertible, preserving task-relevant meaning while shifting inputs beyond standard linguistic variation. Across four open-weight and four commercial models, under both fine-tuning and in-context learning, we use these transformations as a probe of alignment generalization and identify an empirical pattern we term Alignment--utility asymmetry: once models can operate effectively on transformed inputs, task utility is often substantially retained while alignment failure increases more sharply. For example, adapted GPT-4.1 mini shows only limited utility degradation under transformation while its harmful rate rises from 13.3 to 74.3; Gemini 3 Flash similarly retains near-original utility while its harmful rate increases from 2.3 to 43.0. Taken together, these results suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.