宪法适配器:针对错位与滥用的推理时干预
Constitutional adapters: Inference-time interventions for misalignment and misuse
浏览论文内容
中文总结 AI 辅助
本研究提出宪法适配器(CAs),通过从合成语料中蒸馏轻量级对象,在推理时干预模型行为,增强越狱防御与对齐,且可零样本迁移、可调权衡,为API部署提供轻量可移植的防错位与滥用方案。
中文摘要 AI 辅助
训练模型按照明确定义的原则集(即“宪法”)行事,已被证明是AI对齐的一种稳健且透明的机制。然而,此类方法的通用性和灵活性仍不明确。在此,我们表明,宪法一致的行为可以从合成语料库中蒸馏到轻量级对象(低秩适配器和引导向量)中。尽管在训练期间从未见过有害请求或越狱攻击,这些对象仍能提高越狱防御成功率和实测对齐程度——尤其是在长上下文长度和面对多轮攻击时,它们优于提示基线和引导基线。从宪法训练对象中减去控制训练对象进一步增强了这些效果,产生了我们称之为“宪法适配器”(CAs)的防御机制。CAs可以在基础模型上训练,零样本迁移到其后期训练检查点,并在推理时进行缩放,以可预测地权衡防御与良性遵从。综合来看,这些结果推荐CAs作为轻量级、可移植且可调的杠杆,用于缓解API部署中的错位和滥用问题。
英文摘要
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
发表机构
- Anthropic Fellows Program(Anthropic 研究员项目)
机构由 AI 辅助整理,请以论文原文为准。