arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁搭建安全桥梁?识别并靶向跨语言共享安全路径

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

Shuyi Miao, Wangjie Qiu, Pengyang Shao, Canran Xiao, Fei Shen, Zhiming Zheng, Tat-Seng Chua

arXiv 2608.09095首次发表:更新:

AI 中文总结

本研究识别出跨语言共享安全路径,提出基于该路径的靶向对齐方法,仅更新少量参数即可提升非高资源语言安全性并保留模型通用能力。

AI 中文摘要

揭示大型语言模型(LLM)安全能力背后的内部机制,对开发可信人工智能至关重要。当前,多语言安全的机制可解释性研究大多局限于孤立神经元等局部组件,这种静态且碎片化的视角忽略了组件间的协同作用,无法阐明安全信号如何在模型内部动态传播以最终驱动安全决策。本研究突破孤立神经元的局限,识别并靶向安全信号传播过程中形成的跨层功能路径,从而揭示驱动跨语言安全差距的机制。具体而言,我们首先识别单语言安全路径并验证其对拒绝有害请求的影响;后续跨语言分析揭示了跨语言共享安全路径的稀疏子集,证实该交集是将安全能力从高资源(HR)语言转移到非高资源(NHR)语言的内部桥梁。基于这些机制发现,我们提出一种基于跨语言共享安全路径的路径靶向对齐方法。实验结果表明,仅更新小部分路径参数即可显著提升NHR语言的安全性,同时在很大程度上保留模型的通用能力。

英文摘要

Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this static and fragmented perspective overlooks the synergy among components and fails to elucidate how safety signals dynamically propagate within the model to drive safety decisions ultimately. In this work, we move beyond isolated neurons to identify and target the cross-layer functional pathways formed during safety signal propagation, thereby uncovering the mechanisms driving the cross-lingual safety gap. Specifically, we first identify monolingual safety pathways and validate their impact on refusing harmful requests. Subsequent cross-lingual analyses reveal a sparse subset of cross-lingual shared safety pathways, confirming that this intersection acts as the internal bridge transferring safety capabilities from high-resource (HR) languages to non-high-resource (NHR) languages. Building on these mechanistic findings, we propose a pathways-targeted alignment method based on the cross-lingual shared safety pathways. Experimental results show that updating only a small fraction of pathway parameters significantly improves safety in NHR languages while largely preserving the model's general capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑