arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.07918cs.LGcs.AIcs.CLcs.CR

通过潜在人格特质实现语言模型的高效安全对齐

Efficient Safety Alignment of Language Models via Latent Personality Traits

Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型安全方法易受攻击问题,提出潜在人格对齐(LPA)方法,基于66条陈述训练,通过假设人格表征与避害共享结构,实现高攻击成功率、轻量级训练及良好泛化性。

中文摘要 AI 辅助

当前大语言模型的安全方法易受对抗攻击,促使人们研究更强大的替代方法。潜在对抗训练(LAT)是最有效的防御方法之一,但会降低实用性且需要在大量有害提示数据集上进行训练。我们引入潜在人格对齐(LPA),它仅基于从心理测量人格文献中提取的66条与伤害无关的陈述进行对抗训练,取代了明确的伤害拒绝。我们假设基于人格的表征与避免伤害共享潜在结构,因此对抗性地稳定它们会隐式地约束越狱攻击所利用的子空间。LPA在HarmBench上通过直接请求和五种越狱方法实现了近乎零的攻击成功率,在训练期间从未见过有害内容,并且在标准基准测试中没有性能损失。此外,训练过程轻量级,整个过程在单个GPU上只需几分钟,使用的示例比标准LAT少75倍。广泛的消融实验证明了我们方法的鲁棒性、效率和泛化性。

英文摘要

Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75x fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method.

发表机构

  • Mila, Quebec AI Institute(米拉,魁北克人工智能研究所)
  • McGill University(麦吉尔大学)
  • LawZero
  • Université de Montréal(蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑