arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27081cs.AIcs.CLcs.CRcs.LG

面向大语言模型安全的在线策略蒸馏:一种针对模板鲁棒性的重对齐路由方法

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Yongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao, Sheng Wen

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有LLM安全防御的模板鲁棒性不足等问题,提出ROPD框架,通过建模输出分布差异实现重对齐,大幅降低模板不匹配风险,在防御有效性和能力保留上优于基线方法,建立了新的鲁棒重对齐标准。

中文摘要 AI 辅助

微调是使大语言模型(LLM)专业化的主流范式,但它暴露了一个关键漏洞:恶意数据提供者可在下游语料库中嵌入有害行为,使模型在保留专业技能的同时可按需违背人类价值观。现有安全重对齐防御措施在实践中常因三大关键局限失效:它们频繁导致专业技能的灾难性遗忘;当防御者无法观察攻击者的提示模板时,其有效性会崩溃;成功重对齐的模型仍易通过简单的系统提示切换被重新越狱。为应对这些挑战,我们提出基于路由的在线策略蒸馏(Routing-based On-Policy Distillation,ROPD),这是一种新型重对齐框架,用于建模对齐与受损输出概率分布之间的差异,而非拟合特定提示模板。我们开展了大量实验,在三个数据集、三个具有不同对齐强度的基础模型上,将ROPD与四种最先进的基线进行比较。结果表明,当基线防御措施面临模板不匹配时,通常会伴随下游任务性能的严重下降;相比之下,ROPD大幅降低了模板不匹配风险,在防御有效性和能力保留方面均保持了卓越的鲁棒性。尽管我们的分析显示ROPD并非完全不受模板偏移影响,但其性能下降与现有方法相比可忽略不计,为鲁棒的LLM重对齐建立了新的标准。

英文摘要

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susceptible to re-jailbreaking via simple system prompt switches. To address these challenges, we propose Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates. We conduct extensive experiments comparing ROPD against four state-of-the-art baselines across three datasets and three base models with varying alignment strengths. Our results demonstrate that when baseline defenses face template mismatches, often accompanied by severe degradation in downstream task performance. In contrast, ROPD substantially mitigates template-mismatch risks, maintaining superior robustness in both defense effectiveness and capability preservation. While our analysis indicates ROPD is not entirely immune to template shifts, its performance degradation is negligible compared to existing methods, establishing a new standard for robust LLM realignment.

发表机构

  • Tsinghua University(清华大学)
  • Swinburne University of Technology(斯威本科技大学)
  • EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑