arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言安全信号是多层次的:过滤降低安全性的数据以构建更安全的LLM

Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs

Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan

arXiv 2609.22144首次发表:更新:

发表机构

Beihang University; Zhongguancun Laboratory; Hangzhou Innovation Institute, Beihang University; Qingdao Research Institute, Beihang University(北京航空航天大学; 中关村实验室; 北京航空航天大学杭州创新研究院; 北京航空航天大学青岛研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多语言模型微调中安全数据过滤问题,提出多层框架MMSAFE,捕获跨语言共享及语言特定安全信号,将平均有害响应率降低60%,优于单层基线。

AI 中文摘要

在大语言模型微调过程中保持安全对齐至关重要,然而,近期研究表明,即使是良性的微调数据也可能包含悄然削弱安全对齐的降低安全性的样本。现有方法通常利用来自单个安全敏感层的表示来识别此类样本。尽管这一假设在单语言环境中已显示出有效性,但由于跨语言表示模式的潜在差异,其对于多语言模型的有效性仍不明确。通过跨语言分析,我们发现敏感层在不同语言间仅部分共享,安全相关信号往往分布在多个层中。受这些观察启发,我们提出了MMSAFE,一个用于多语言降低安全性数据识别的多层框架,该框架同时捕获共享的和语言特定的安全信号。在多个模型、语言和安全基准上的大量实验表明,与随机过滤相比,MMSAFE将平均有害响应率降低了60%,并且比最强的单层基线取得了更强的平均性能,证明了多层建模对于稳健的多语言安全对齐的有效性。

英文摘要

Preserving safety alignment during large language models fine-tuning is critical, however, recent studies have demonstrated that even benign fine-tuning data may contain safety-degrading samples that silently undermine safety alignment. Existing approaches typically identify such samples using representations from a single safety-sensitive layer. While this assumption has shown effectiveness in monolingual settings, its validity for multilingual models remains unclear due to potential cross-lingual differences in representation patterns. Through a cross-lingual analysis, we show that sensitive layers are only partially shared across languages, with safety-relevant signals often distributed across multiple layers. Motivated by these observations, we propose MMSAFE, a multi-layer framework for multilingual safety-degrading data identification that captures both shared and language-specific safety signals. Extensive experiments across multiple models, languages, and safety benchmarks demonstrate that MMSAFE reduces the average harmful-response ratio by 60% compared with random filtering and achieves stronger average performance than the strongest single-layer baseline, demonstrating the effectiveness of multi-layer modeling for robust multilingual safety alignment.

CommentsAccepted to EMNLP 2026 Main Conference. 16 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑