发表机构
Macau University of Science and Technology(澳门科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SafeDepth是一个轻量级即插即用框架,通过路由器选择性地跳过Transformer层并配合适配器,在保持任务性能的同时减少计算量并降低有害响应率,实验表明其在多个基准上优于FlexiDepth。
AI 中文摘要
近期研究表明,并非每个令牌都需要经过所有Transformer层,这促使了令牌级自适应模型的产生,该类模型可选择性地跳过某些层以减少计算量。我们的实验表明,这些执行选择也会影响安全性:现有的令牌级自适应推理模型相比其参考骨干模型,表现出更高的有害响应率。安全影响取决于跳过了哪些层,以及跳过发生在提示处理阶段还是答案生成阶段。我们进一步发现,识别有害请求和生成拒绝响应依赖于不同的逐层计算。我们引入了SafeDepth,一个轻量级、即插即用的框架,利用选择性层执行来同时提升安全性和效率。SafeDepth学习保留支持安全响应的计算,并绕过那些导致不安全生成的计算。一个路由器根据令牌表示、层位置和推理阶段来选择层执行,同时一个适配器支持被跳过的路径。我们联合训练这些模块,以减少计算量和有害响应,同时保持良性任务上的性能,旨在实现更好的安全-效率权衡。预训练的骨干模型全程保持冻结,无需额外的预训练。在Llama-3-8B-Instruct上的实验结果表明,SafeDepth相对于全深度推理减少了计算量,同时基本保持了任务性能。与FlexiDepth相比,它在所有五个有害请求基准上降低了不安全响应率,其中在HarmBench-HJ上绝对降幅最大(从55.44%降至25.25%),同时将XSTest上的错误拒绝率从12.0%降至4.0%。
英文摘要
Recent studies suggest that not every token needs to pass through all Transformer layers, motivating token-level adaptive models that selectively skip layers to reduce computation. Our experiments show that these execution choices also affect safety: existing token-level adaptive reasoning models exhibit higher harmful-response rates than their reference backbones. The safety effects depend on which layers are skipped and whether skipping occurs during prompt processing or answer generation. We further find that recognizing harmful requests and producing refusals depend on different layer-wise computations. We introduce SafeDepth, a lightweight, plug-in framework that uses selective layer execution to improve both safety and efficiency. SafeDepth learns to retain computations that support safe responses and bypass those that contribute to unsafe generation. A router selects layer execution based on token representations, layer position, and inference phase, while an adapter supports the skipped paths. We jointly train these modules to reduce computation and harmful responses while preserving performance on benign tasks, targeting a better safety-efficiency trade-off. The pretrained backbone remains frozen throughout, requiring no additional pretraining. Experimental results on Llama-3-8B-Instruct show that SafeDepth reduces computation relative to full-depth inference while largely preserving task performance. Compared with FlexiDepth, it lowers unsafe-response rates across all five harmful-request benchmarks, with the largest absolute decrease on HarmBench-HJ (from 55.44% to 25.25%), while reducing the XSTest false-refusal rate from 12.0% to 4.0%.