发表机构
Google Research; Cambridge University; Google DeepMind; Google; Tel Aviv University(谷歌研究院; 剑桥大学; 谷歌DeepMind; 谷歌; 特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过谱优化调制Transformer模型中安全指令的特征值,提出对比安全损失以平衡有害查询上的强调与无害查询上的抑制,实现帕累托改进的安全指令。
AI 中文摘要
在基于Transformer的语言模型中,上下文令牌可以被吸收到模型权重中,形成一个乘法算子。我们在安全指令的背景下研究该算子,并表明影响其主特征值可以调制指令对生成的塑造强度。我们推导出一个对比安全损失,其中包含一个抑制权重,用于控制在有害查询上强调安全指令与在无害查询上抑制该指令之间的权衡。改变抑制权重映射出攻击成功率与过度拒绝率之间的关系,支持了该算子特征值作为指令影响力连续调节旋钮的假设。此外,这种关系相对独立于安全损失的参数化方式,对于适当的抑制权重值,能够产生帕累托改进的安全指令。
英文摘要
Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.