HiRoute:用于大语言模型安全对齐的分层路由提示调优
HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models
中文总结 AI 辅助
HiRoute是一种输入自适应分层提示调优框架,通过分层路由器与多粒度提示专家,提升大语言模型安全对齐效果,在多个安全基准上兼顾安全性与响应有用性,减少过度拒绝。
中文摘要 AI 辅助
大语言模型(LLMs)仍易受到有害请求和越狱攻击的影响。基于提示调优的参数高效安全对齐方法通常依赖单个全局提示或外部选择的提示模块,这种静态设计难以在维持跨类别安全边界的同时,生成针对特定风险的建设性响应,并避免对良性输入的过度拒绝。为解决这些局限,我们提出HiRoute,一种输入自适应的分层提示调优框架,它将类别无关的安全控制与类别特定的响应引导分离开来。HiRoute首先在冻结的LLM提取的表示上训练一个轻量级分层路由器,以联合检测有害意图并预测多标签风险分数;随后冻结骨干模型和路由器,使用交替梯度更新的偏好优化来学习共享的粗粒度提示和一组细粒度提示专家作为连续嵌入。推理时,良性输入绕过安全分支,而风险输入则使用共享提示与路由器加权的风险特定提示专家混合进行处理。在三个指令调优模型上的实验表明,HiRoute在多个安全基准上实现了高安全率,同时保留了安全响应的有用性,减少了过度拒绝,并在通用任务上保持了有竞争力的性能。
英文摘要
Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, we propose HiRoute, an input-adaptive hierarchical prompt-tuning framework that separates category-agnostic safety control from category-specific response guidance. HiRoute first trains a lightweight hierarchical router on representations extracted from a frozen LLM to jointly detect harmful intent and predict multi-label risk scores. It then freezes both the backbone model and the router and uses preference optimization with alternating gradient updates to learn a shared coarse-grained prompt and a set of fine-grained prompt experts as continuous embeddings. At inference time, benign inputs bypass the safety branch, whereas risky inputs are processed using the shared prompt together with a router-weighted mixture of risk-specific prompt experts. Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.