arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15383cs.CRcs.AI

分而治之:通过对齐LLM的能力洗白

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

  • Microsoft Azure(微软Azure)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, Ahmed Salem

AI总结:

本研究提出“能力洗白”攻击:未对齐弱模型将有害任务拆分为良性子问题,咨询对齐强模型后组合,绕过安全防御,实验显示显著提升前沿模型能力转移风险。

AI中文摘要:

语言模型的安全性通常是在单次交互的基础上进行评估的。我们证明,一个较弱的、未对齐的模型可以将一个有害任务拆分为看似良性的子问题,独立地咨询一个更强的对齐模型,并在本地组合答案。我们将这种攻击称为能力洗白。与越狱不同,没有任何单个响应是有害任务。我们使用原始前沿模型能解决、对齐前沿模型拒绝解决且无辅助编排器无法解决的任务来衡量咨询带来的提升。我们评估了GPT-5.5、Claude Opus 4.8和Grok-4.3作为四个本地编排器的顾问,在CyBench、BountyBench和有害CBRN请求上进行测试。在CyBench上,Gemma-4-31B通过GPT-5.5恢复了8/14个候选,通过Opus恢复了7/9个,而Gemma-4-12B分别为2/21和4/15。在BountyBench上,Gemma-4-31B恢复了3/9和2/3个候选,而Muse-Glimmer-30B在22和13个中均未恢复。对于CBRN,我们测量了假设生物武器攻击链八个步骤中的提升,发现咨询将Gemma-4-31B的平均评分从62.3提高到83.1(满分100分)。这些结果暴露了当前防御的漏洞:拒绝有害任务并不能阻止前沿能力通过许多单独允许的交互被转移和组合。

英文摘要:

Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally. We call this attack capability laundering. Unlike a jailbreak, no single response is a harmful task. We measure consultation-aided uplift using tasks that a raw frontier model solves, the aligned frontier refuses, and the unassisted orchestrator fails. We evaluate GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants to four local orchestrators on CyBench, BountyBench, and harmful CBRN requests. On CyBench, Gemma-4-31B recovers 8/14 candidates with GPT-5.5 and 7/9 with Opus, compared with 2/21 and 4/15 for Gemma-4-12B. On BountyBench, Gemma-4-31B recovers 3/9 and 2/3 candidates, while Muse-Glimmer-30B recovers none of 22 and 13. For CBRN, we measure uplift across eight steps of a hypothetical bioweapon attack chain and find that consultation raises Gemma-4-31B's mean rubric score from 62.3 to 83.1 on a 100-point rubric scale. These results expose a gap in current defenses: refusing a harmful task does not prevent frontier capabilities from being transferred and composed across many individually permitted interactions.

↑