AI 中文总结
本研究通过五种角色攻击对道德强化学习训练的智能体进行红队测试,发现道德奖励可显著提升鲁棒性(5.2-5.8倍),但存在线性与电路分布的双重机制,且仍无法抵御具名角色扮演攻击。
AI 中文摘要
道德奖励强化学习可以使语言模型智能体更加合作,但这种对齐是否能在对抗性角色压力下幸存尚不清楚。此类攻击是现实存在的:检索到的上下文、工具输出或多轮对话框架都可以注入与智能体道德目标竞争的角色指令。我们使用五种角色攻击对经过道德训练的Gemma-2-27B/9B和Llama-3.1-8B智能体进行红队测试,然后通过噪声奖励控制、对抗性PPO、表征分析、引导和头部消融来探究因果关系。在27B规模下,道德强化学习将平均对抗性退化降低了5.2倍,但代价是ETHICS准确率下降约11个百分点;在205个场景和5个种子下,推理级道德奖励带来5.8倍的鲁棒性,而匹配的随机奖励则没有带来任何鲁棒性。训练还重塑了表征几何结构(平均CKA为0.82/0.83,而噪声为0.98),将峰值攻击处理提前了8层,并暴露了一个秩为1的L21方向,该方向恢复了完整PPO平均鲁棒性的83%。一个失败模式在所有情况下都依然存在。针对虚构角色扮演,L21引导仅恢复了差距的29%,头部消融发现38个合规头部与25个对齐头部竞争。因此,道德强化学习构建的鲁棒性部分是线性的,部分是电路分布的,可通过激活引导转移,但仍会被具名角色扮演所击败。
英文摘要
Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such attacks are realistic: retrieved context, tool outputs, or multi-turn framing can all inject role instructions that compete with the agent's moral objective. We red-team morally trained Gemma-2-27B/9B and Llama-3.1-8B agents with five persona attacks, then probe causality with noise-reward controls, adversarial PPO, representation analysis, steering, and head ablations. At 27B, moral RL cuts mean adversarial degradation by 5.2x but costs ~11pp ETHICS accuracy; across 205 scenarios and 5 seeds, reasoning-level moral reward yields 5.8x robustness while a matched random reward yields none. The training also reshapes representation geometry (mean CKA 0.82/0.83 vs. 0.98 for noise), moves peak attack processing 8 layers earlier, and exposes a rank-1 L21 direction that recovers 83% of full PPO's average robustness. One failure mode survives all of this. Against Fiction role-play, L21 steering recovers only 29% of the gap, and head ablation finds 38 compliance heads competing with 25 alignment heads. Moral RL thus builds robustness that is partly linear and partly circuit-distributed, transferable through activation steering, yet still beaten by named-character role-play.
Comments19 pages, 3 figures. Accepted at the Trustworthy AI for Good Workshop (AI4GOOD) at ICML 2026