arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DataRx:面向更安全的大语言模型任务特定微调的缺失感知采样方法

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

Junbo Zhang, Qianli Zhou, Xinyang Deng, Wen Jiang

arXiv 2608.04322首次发表:更新:

发表机构

Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DataRx缺失感知采样方法,通过高维隐表示量化安全信号差距,仅用1% BeaverTails安全样本,即可大幅降低Llama3-8B-Instruct的攻击成功率,还可与现有安全数据合成方法结合提升防御效果。

AI 中文摘要

任务特定微调可提升大语言模型(LLM)在下游任务上的性能,但本研究发现其也会削弱对齐后LLM的安全护栏。微调过程中,融合安全数据是常用的安全保留策略,尽管此前研究表明随机混合安全数据可缓解安全性能下降,但仍不清楚为何部分安全示例更有效。本文提出DataRx,一种用于选择安全关键示例的缺失感知采样方法,其基于如下假设:当所选安全示例提供的安全信号填补了LLM安全能力的缺失部分时,该安全样本更有效。DataRx的核心见解是利用高维隐表示而非离散token,量化目标模型原生响应与安全参考响应之间的安全信号差距。结果显示,仅需额外添加1%来自BeaverTails的安全样本,DataRx即可将Llama3-8B-Instruct在7项下游任务上的平均攻击成功率,从随机采样下的59.23%降至13.70%。此外,DataRx可与现有安全数据合成方法结合,进一步增强微调期间的安全防御,希望DataRx能推动更多以数据为中心的防御研究。

英文摘要

Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific fine-tuning can also weaken the safety guardrails of aligned LLMs. A widely adopted strategy for preserving safety during fine-tuning is to incorporate safety data. Although previous studies have shown that randomly mixing safety data can alleviate safety degradation, the underlying principle determining why some safety examples are more effective than others still remains unclear. In this paper, we propose DataRx, a missingness-aware sampling method for selecting safety-critical examples. DataRx is based on the hypothesis that a safety sample is more effective when the selected examples provide safety signals that fill the missing parts of LLMs' safety capabilities. DataRx's key insight is leveraging high-dimensional hidden representations rather than discrete tokens to quantify the safety signal gap between the target model's native response and the safety reference response. The results show that, with only 1% additional safety samples from BeaverTails, DataRx reduces the average attack success rate of Llama3-8B-Instruct across seven downstream tasks from 59.23% under random sampling to 13.70%. In addition, DataRx can be combined with the existing safety data synthesis method to further enhance safety defenses during fine-tuning. We hope that DataRx will inspire more data-centric defense research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑