发表机构
Nanjing University of Science and Technology; The University of Sydney; National University of Singapore; University of New South Wales; Nanjing University(南京理工大学; 悉尼大学; 新加坡国立大学; 新南威尔士大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型基础模型易受路由级白盒攻击的问题,提出动态路由自适应对齐框架,通过引入动态补偿路由提升模型安全稳健性且保留通用实用性。
AI 中文摘要
随着大型基础模型(LFMs)在开放环境中的广泛部署,安全威胁正从黑盒越狱转向可直接识别并破坏内部安全神经元或路由的白盒攻击。然而,现有的安全防御通常依赖静态安全单元或固定拒绝路径,导致模型极易受到针对性的路由级白盒攻击。为此,我们提出动态路由自适应对齐(DRAA)框架,该框架引入动态补偿路由,以在安全路由被破坏时保持稳健的拒绝行为。具体而言,我们首先通过对比安全与不安全校准样本的内部激活来识别并定位模型的安全路由;随后,DRAA会屏蔽该安全路由以诱导因果失败案例,并选择性挖掘由此产生的防御失败,从而构建感知失败的偏好对。大量实验表明,DRAA可有效重构模型安全的底层路径依赖,大幅提升对路由级白盒攻击的稳健性,同时保留通用实用性。
英文摘要
With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes. However, existing safety defenses often rely on static safety units or fixed refusal pathways, leaving models highly vulnerable to targeted route-level white-box attacks. For that, we propose dynamic routing adaptive alignment (DRAA), a framework that introduces dynamic compensatory routes to preserve robust refusal behavior when the safety route is compromised. Specifically, we first identify and localize the model's safety route by contrasting internal activations between safe and unsafe calibration samples. DRAA then masks this safety route to induce causal failure cases and selectively mines the resulting defense failures, thereby constructing failure-aware preference pairs. Extensive experiments demonstrate that DRAA effectively restructures the underlying pathway dependence of model safety, substantially improving robustness against route-level white-box attacks, while preserving general utility.