AI 中文总结
本研究提出能力路由守卫(CRG),一种针对闭源大型推理模型的模型无关推理时护栏,可有效缓解多种以推理为中心的越狱攻击,同时保留良性效用并避免过度拒绝问题。
AI 中文摘要
大型推理模型(LRMs)暴露出一种新的安全故障模式:对抗性提示词可操纵推理上下文、任务分解或能力解释,从而将有害目标处理为合法的推理步骤。现有防御措施包括安全提醒、外部分类器和自检查包装器,通常较为脆弱,因为它们要么直接检查对抗性提示词,要么要求目标模型在攻击所利用的同一层面执行额外的安全推理。我们提出能力路由守卫(Capability-Routed Guard,CRG),这是一种针对闭源LRMs的模型无关推理时护栏,防御者无法检查隐藏的推理痕迹或修改模型权重。CRG将提示词防御重新表述为能力路由问题:一个侧通道控制器首先构建用户授权任务、活跃上下文、安全证据和能力转移风险的可信表示,将可执行意图与不可信的推理上下文分离。该表示支持特定路由执行,使CRG能够阻止高风险请求、约束模糊请求,并通过可信活跃上下文转发低风险请求。最后,CRG应用TraceCheck验证与授权任务的一致性,并调用受限回退以保留低风险良性提示词的效用。大量实验表明,CRG在保留良性效用的同时有效缓解了多种以推理为中心的越狱攻击,且避免了常见的过度拒绝问题。进一步分析显示其组件贡献了互补的益处,凸显了协调防御机制对保障大型推理模型安全的重要性。
英文摘要
Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmful objectives are processed as legitimate reasoning steps. Existing safeguards, including safety reminders, external classifiers, and self-checking wrappers, are often brittle because they either inspect the adversarial prompt directly or ask the target model to perform additional safety reasoning on the same surface that attacks exploit. We introduce Capability-Routed Guard (CRG), a model-agnostic inference-time guardrail for closed-source LRMs, where defenders cannot inspect hidden reasoning traces or modify model weights. CRG reframes prompt defense as a capability-routing problem: a side-channel controller first constructs a trusted representation of the user's authorized task, active context, safety evidence, and capability-transfer risk, separating executable intent from untrusted reasoning context. This representation supports route-specific execution, allowing CRG to block high-risk requests, constrain ambiguous ones, and forward low-risk requests through trusted active context. Finally, CRG applies TraceCheck to verify consistency with the authorized task and invokes a restricted fallback to preserve utility for low-risk benign prompts. Extensive experiments demonstrate that CRG effectively mitigates diverse reasoning-centric jailbreaks while preserving benign utility and avoiding common over-refusal issues. Further analysis shows that its components contribute complementary benefits, highlighting the importance of coordinated defense mechanisms for securing large reasoning models.