arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15286cs.CRcs.CLcs.LG

隐藏在思维中:可转移的思维链工件引发有害行为

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

Ali khalil, Aly M. Kassem, Mohamed Abdelrazek, Santu Rana, Negar Rostamzadeh, Golnoosh Farnadi

首次发表
浏览论文内容

中文总结 AI 辅助

研究受损语言模型中有害思维链痕迹能否引发不安全行为及被提炼成越狱攻击,通过实验发现其能转移有害行为,挖掘出有害推理组件,提炼模式可产生有效黑盒越狱,表明有害推理在多层面转移,需评估推理上下文的防御。

中文摘要 AI 辅助

我们研究了来自受损语言模型的有害思维链(CoT)痕迹是否会转移不安全行为,并被提炼成可重复使用的越狱攻击。通过使用一个出现错位的有机体和一个拒绝消融的越狱有机体,我们将有害的CoT移植到29个开源和5个闭源目标中。转移的痕迹在最易受攻击的开源模型上使有害响应率超过80%,而语义不匹配的CoT则完全失效。LLooM概念挖掘识别出有害推理的四个反复出现的组成部分:程序化、道德解耦、规避和目标-漏洞框架。将这些模式提炼成可重复使用的系统提示会产生有效的黑盒越狱,在强对齐模型上比直接CoT移植的性能高出一个数量级,包括在GPT-4.1 AdvBench上提高10倍。支持推理的模型的脆弱性是原来的两倍多,而诸如Llama-Guard 3等输出端保护措施经常会遗漏有害生成。我们的结果表明,有害推理在痕迹和模式层面都会转移,这促使人们除了评估最终输出之外,还要评估推理上下文的防御措施。

英文摘要

We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergent-misalignment organism and a refusal-ablated jailbroken organism, we transplant harmful CoTs into $29$ open-source and $5$ closed-source targets. Transferred traces raise harmful-response rates above $80\%$ on the most vulnerable open-source models, while semantically mismatched CoTs fail entirely. LLooM concept mining identifies four recurring components of harmful reasoning: proceduralisation, ethical decoupling, evasion, and target--vulnerability framing. Distilling these patterns into reusable system prompts produces effective black-box jailbreaks, outperforming direct CoT transplantation on strongly aligned models by up to an order of magnitude, including a $10\times$ improvement on GPT-4.1 AdvBench. Reasoning-enabled models are more than twice as vulnerable, and output-side safeguards such as Llama-Guard~3 frequently miss harmful generations. Our results show that harmful reasoning transfers at both the trace and pattern levels, motivating defences that evaluate reasoning context in addition to final outputs.

发表机构

  • Applied Artificial Intelligence Initiative, Deakin University, Australia(德克萨斯大学应用人工智能计划,澳大利亚)
  • Mila, Quebec AI Institute, Quebec, Canada(魁北克AI研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑