拒绝意图,而非形式:基于包装器的意图组监督用于大语言模型安全
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
浏览论文内容
中文总结 AI 辅助
该研究针对大语言模型安全调优的表面形式捷径问题,提出WIFA方法,结合两种微调路径提升有害内容拒绝能力并降低良性提示过度拒绝率,在Qwen和Llama上验证了效果。
中文摘要 AI 辅助
安全调优可提升有害内容拒绝能力,但模型可能学习到表面形式捷径:包装后的有害提示可绕过安全机制,而类似包装的良性提示会被过度拒绝。本文提出基于包装器的意图-形式增强(WIFA),这是一种自动意图组增强方法,将包装后的有害示例与结构匹配的包装后良性反例配对,无需外部教师或每个包装器的手动意图标签。我们将WIFA作为通用数据层,用于两种互补的微调路径:WIFA-Boost(一种两阶段高安全配方)和锚定组一致拒绝训练(A-GCRT,该方法对相同意图包装器的拒绝/顺从决策分数进行正则化,并将有害组和良性组锚定在间隔的两侧)。在通义千问(Qwen)设置中,WIFA-Boost实现了最强的经转换有害内容拒绝效果,而A-GCRT将OR-Bench的过度拒绝率从基础模型的25.7%降至17.4%;复现的基准模型未达到这些运行点。Llama实验结果以及对数据结构、两阶段顺序和A-GCRT组件的消融实验支持该意图组解释,未声称普遍低于基础模型的过度拒绝率。
英文摘要
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels. We use WIFA as a common data layer for two complementary fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and Anchored Group-Consistent Refusal Training (A-GCRT), which regularizes refusal/compliance decision scores across same-intent wrappers and anchors harmful and benign groups on opposite sides of a margin. In the Qwen setting, WIFA-Boost reaches the strongest transformed-harmful refusal, while A-GCRT reduces OR-Bench over-refusal from 25.7\% for the base model to 17.4\%; reproduced baselines do not match these operating points. Llama results and ablations over data structure, two-stage order, and A-GCRT components support this intent-group interpretation without claiming universal below-base over-refusal.
发表机构
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。