通过反对齐少样本对话暴露实现大型推理模型的对齐
Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure
- Wayne State University(韦恩州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SRCF攻击,利用反对齐少样本对话引导大型推理模型推理,导致不安全生成或错误拒绝;并设计ARCF防御,通过暴露反对齐上下文并强制对齐目标,在保持实用性的同时提升安全性和有用性。
AI中文摘要:
大型推理模型(LRMs)依赖显式的思维链(CoT)推理和大上下文窗口,在复杂任务上取得强性能,但这些特性也引入了新的攻击面。我们证明,通过在提示前添加包含显式CoT轨迹的反对齐少样本对话,可以系统地引导LRMs的推理过程,导致对有害查询产生不安全生成,对良性查询产生无根据的拒绝。我们将这种攻击形式化为SRCF(通过反对齐少样本对话引导推理),它仅通过灵活的对话界面运作,无需访问模型的参数和梯度。我们的关键洞见是,SRCF利用了一种对抗性泛化问题,引发表征漂移,使良性输入和有害输入的表征向相似方向移动。这一观察启发了我们的训练后防御方法ARCF(通过反对齐少样本对话对齐推理),该方法在强制对齐目标的同时,将模型暴露于反对齐对话上下文中。ARCF与现有训练后方法兼容,在不降低实用性的情况下持续提升安全性和有用性。
英文摘要:
Large Reasoning Models (LRMs) rely on explicit chain-of-thought (CoT) reasoning and large context windows to achieve strong performance on complex tasks, but these features also introduce new attack surfaces. We show that LRMs' reasoning processes can be systematically steered by prepending counter-aligned few-shot conversations containing explicit CoT traces, leading to unsafe generations on harmful queries and unwarranted refusals on benign ones. We formalize this attack as SRCF (Steering Reasoning via Counter-Aligned Few-shot Conversations) that operates solely through a flexible conversational interface and requires no access to the model's parameters and gradients. Our key insight is that SRCF exploits an adversarial generalization issue that induces a representation drift, causing the representations of benign and harmful inputs to shift in a similar direction. This observation motivates our post-training defense, ARCF (Aligning Reasoning via Counter-Aligned Few-Shot Conversations), which exposes models to counter-aligned conversational contexts while enforcing aligned targets. ARCF is compatible with existing post-training methods and consistently improves safety and helpfulness without degrading utility.