actr:对齐推理大语言模型中的思维与响应以实现多语言安全
ACTR: Aligning Thoughts and Responses for Multilingual Safety in Reasoning LLMs
浏览论文内容
中文总结 AI 辅助
针对推理大语言模型在多语言下的安全漏洞,提出ACTR框架,通过思维差距分数和神经元选择优化对齐思维与响应,降低越狱攻击成功率并保持性能。
中文摘要 AI 辅助
确保推理大语言模型(LLM)在不同语言中的安全性对于其可靠部署至关重要。然而,当这些模型在非高资源语言中遭受越狱攻击时,即使其推理轨迹识别出安全风险,它们仍可能生成不安全的响应。为解决此问题,我们提出对齐跨语言思维与响应(ACTR),这是一个通过加强现有安全推理的利用来改善多语言安全对齐的框架。具体而言,我们首先提出思维差距分数(TGS),用于比较响应生成过程中不同语言下推理轨迹对注意力输出的归一化贡献,并使用推理轨迹替换来度量跨语言安全差距。接下来,利用越狱查询语料库,我们通过神经元掩蔽引起的响应表示变化来评估神经元重要性,并比较启用和禁用推理时获得的高重要性神经元集,以识别支持安全推理使用的安全思维神经元。最后,我们设计了神经元选择性一致性优化(NSCO),它使用冻结的评判模型来奖励推理轨迹与响应的安全类别之间的一致性,同时仅更新与所选神经元相关的参数,无需人工标注的响应或偏好数据。在两个推理模型上,ACTR在AdvBench-X和MultiJail上实现了比评估的最先进方法更低的平均攻击成功率,安全增益扩展到未见语言,同时保持或提高了多语言知识和数学推理任务的平均性能,并限制了良性请求的误拒。警告:本文包含不安全内容的示例。
英文摘要
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.
发表机构
- Nanjing University of Science and Technology(南京理工大学)
- National University of Singapore(新加坡国立大学)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。