arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示模板在安全对齐知识蒸馏中的作用理解

Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment

Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, Yi Ding

arXiv 2609.30802首次发表:更新:

发表机构

University of Tennessee, Knoxville; Georgia State University(田纳西大学诺克斯维尔分校; 佐治亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨知识蒸馏中提示模板选择对学生模型安全对齐的影响,发现非聊天模板能更好地保留内部表征并减少安全退化。

AI 中文摘要

先前的研究表明,在监督微调(SFT)过程中选择提示模板会显著影响后续安全对齐的稳健性。然而,在从教师模型到学生模型的知识蒸馏(KD)过程中,模板选择的影响在很大程度上仍未得到探索。因此,我们通过分析不同模板配置如何影响学生模型预先存在的安全对齐来填补这一空白。我们观察到,在已对齐的基础指令微调模型中,安全对齐出现了显著退化。具体而言,我们发现使用聊天模板会使模型对有害查询的顺从度高于非聊天模板。这些发现在三个模型上保持一致:LLaMA、Gemma和Qwen模型家族,并在多个安全基准上进行了评估。我们进一步表明,在蒸馏过程中使用非聊天模板能更好地保留基础学生模型的内部表征,而聊天模板蒸馏则会引起更大的表征偏移。代码:此https URL

英文摘要

Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill this gap by analyzing how different template configurations influence the pre-existing safety alignment of the student. We observe a significant degradation of safety alignment present in the aligned base instruct-tuned model. Specifically, we find that utilizing chat templates renders the model more compliant with harmful queries compared to a non-chat template. These findings are consistent across three models: LLaMA, Gemma and Qwen model families and are evaluated across multiple safety benchmarks. We further show that using a non-chat template during distillation better preserves the base student's internal representations, while chat template distillation induces a larger representational shift. Code: https://github.com/anjilab/role-of-prompt-template-in-kd

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑