超越平均安全性:机会约束的大语言模型微调
Beyond Average Safety: Chance-Constrained LLM Fine-tuning
浏览论文内容
中文总结 AI 辅助
本文提出机会约束的大语言模型微调方法,限制安全示例退化比例,通过可微上界和约束感知梯度下降实现尾部感知的安全校正,实验证明其优于现有基线。
中文摘要 AI 辅助
在新目标上微调大型语言模型可以提高有用性、指令遵循或特定领域的性能,但也可能引发安全关键提示上的性能退化。现有的安全保持微调方法通常控制平均安全损失或使用加权辅助惩罚,这可能会掩盖罕见但严重的失败。我们提出了一种用于安全保持微调的机会约束公式,该公式限制了相对于参考模型退化超过规定阈值的安全示例的比例。由于由此产生的经验机会约束包含不连续的指示函数,我们引入了违规率的可微上界,从而得到一个易于处理的保守约束。然后,我们开发了一种约束感知的梯度下降方法,该方法将上界约束视为参数空间中的安全集,并最小限度地修改微调方向以保持可行性。由此产生的更新具有封闭形式,并产生一个尾部感知的安全校正,该校正强调接近或超过退化阈值的示例。我们在三个不同任务和三个模型上对有害微调进行了广泛的实验,结果表明我们的方法始终优于文献中现有的基线。这些结果表明,大语言模型微调中的安全保持更适合被视为可靠性约束优化问题,而不是平均风险正则化。
英文摘要
Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold. Because the resulting empirical chance constraint contains a discontinuous indicator, we introduce a differentiable majorization of the violation rate, yielding a tractable conservative constraint. We then develop a constraint-aware gradient descent method that treats the majorized constraint as a safe set in parameter space and minimally modifies the fine-tuning direction to preserve feasibility. The resulting update admits a closed form and produces a tail-aware safety correction that emphasizes examples near or above the degradation threshold. We conduct an extensive set of experiments on harmful fine-tuning across three different tasks and three models and show that our approach consistently outperforms the baselines that exist in the literature. These results suggest that safety preservation in LLM fine-tuning is better viewed as a reliability-constrained optimization problem than as average-risk regularization.
发表机构
- Johns Hopkins University(约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。