AI 中文总结
针对安全对齐LLM在数据选择后仍易受投毒攻击的问题,提出Bi-QSTO方法,在质量约束下优化有毒样本以穿透选择,实验显示高保留率和攻击有效性。
AI 中文摘要
经过安全对齐的大型语言模型(LLMs)仍然容易受到在少量有害或看似良性的样本上进行微调的影响。然而,先前的研究通常假设有毒样本直接进入下游微调,忽略了实际训练流程中基于质量的选择。为填补这一空白,我们系统地评估了过滤对投毒的抵御效果以及保留数据对下游安全性的影响。结果表明,选择过程移除了许多明显有害的样本,但一些保留的高质量样本仍可能降低模型的安全对齐,这可能是由于它们在逐层梯度水平上具有类似有害的训练更新模式。这些发现共同揭示了一个实际漏洞:降低安全性的影响可以通过保留的高质量样本穿透基于质量的选择。为了检验其系统性可利用性,我们提出了两阶段质量约束的安全退化文本优化(Bi-QSTO),该方法在显式质量约束下优化有毒样本,使其在选择中存活,同时保留其降低安全性的影响。在各种投毒设置、目标模型和过滤率下,Bi-QSTO在选择前后均保持了攻击有效性。即使在90%的过滤率下,以有害种子生成的样本仍实现了超过90%的投毒保留率,以及3.30--4.01的有害分数。其攻击有效性在模型间具有很强的迁移性,其保留优势也泛化到其他选择方法。
英文摘要
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fill this gap, we systematically evaluate both the filtering effects against poisoning and the downstream safety impact of retained data. The results reveal that selection removes many overtly harmful samples, yet some retained high-quality samples can still degrade model safety alignment possibly due to their harmful-like training-update patterns at the layer-wise gradient level. Together, these findings expose a practical vulnerability: safety-degrading influence can pass through quality-based selection via retained high-quality samples. To examine its systematic exploitability, we propose Bi-Stage Quality-Constrained Safety-Degradation Text Optimization (Bi-QSTO), which optimizes poisoned samples under an explicit quality constraint to survive selection while preserving their safety-degrading influence. Across poisoning settings, target models, and filtering rates, Bi-QSTO maintains attack effectiveness before and after selection. Even at 90% filtering, harmful-seeded samples achieve a Poisoning Retention Rate above 90% and Harmful Score of 3.30--4.01. Their attack effectiveness strongly transfers across models and their retention advantage generalizes to additional selection methods.
Comments15 pages, in submission