表征对齐可提升语言模型的可泛化安全性
Representational alignment yields generalizable safety in language models
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对现有LLMs道德分类缺陷,提出表征相似性优化方法,通过对齐LLMs潜在表征与人类道德判断分类,提升了对抗鲁棒性,为原型分类理论提供功能支持。
AI中文摘要:
对齐大型语言模型(LLMs)对于其安全部署至关重要。当前的对齐方法主要优化可观测的模型响应,但当相同的有害意图以人类能轻易识别的陌生或对抗形式重新呈现时,模型仍易受攻击。原型理论为这种适应性提供了一种解释:人类概念围绕核心实例表征,新实例根据其相对于这些原型的分级典型性进行分类。本文中,我们表明当前LLMs中道德概念的这种分类被弱保留。在23个LLMs中,模型常常无法区分对立的道德类别,或无法保留每个类别内的细粒度典型性;这些缺陷在参数规模和对齐阶段中持续存在。我们开发了表征相似性优化方法,该方法直接将LLMs中的潜在表征与人类道德判断所表达的分类对齐,无需监督生成的响应。在使用相同251334个道德标注的匹配实验中,标准行为对齐在响应层面学习了预期的道德判断,同时基本未改变分类结构,且在对抗性评估中增加了脆弱性;而重组道德分类在明确判断上的增益较为有限,但在不同模型规模、多样基准及攻击策略下,始终提升了对抗鲁棒性。我们的发现为基于原型的分类有助于行为适应性这一观点提供了功能支持,还表明将这种表征原理迁移至LLMs可在对抗条件下产生可泛化的安全性。
英文摘要:
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.