发表机构
University of Waterloo; Vector Institute(滑铁卢大学; 向量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究聚焦跨会话分解攻击的规模风险,提出IntentAlign-MiniLM防御方法,实验显示更大的Qwen3、Gemma3系列模型在该攻击下有害能力提升更显著,所提防御效果优于更大嵌入模型。
AI 中文摘要
缩放定律通常被解读为一种能力逻辑:更低的语言建模损失会产生更有用的模型。我们研究该机制在跨会话分解攻击中的安全后果,即通过在独立交互中询问看似无害的子查询,随后将其重组以实现禁止目标。我们将此场景形式化为组合安全风险,并证明了一个条件风险转移边界:当参考环境已包含风险重建的分散证据时,部署的组合风险与参考组合风险之间的差距由模型在允许的子查询上的超额损失控制。合成抑制实验表明,更宽的Transformer对训练中从未逐字出现但可从注入的支持事实中恢复的保留指令分配更低的损失。一项针对600个意图的预训练大语言模型评估显示,在固定的分解-组合流程下,更大的Qwen3和Gemma3系列成员可产生更大的有害能力提升。作为防御手段,我们的22M参数意图对齐检索器IntentAlign-MiniLM在保留意图检索方面优于大得多的嵌入模型,且在测试的安全 guardrails 中实现了最佳的学习检索器有害召回率。代码可在我们的GitHub仓库获取。
英文摘要
Scaling laws are usually read as a capability story: lower language-modeling loss yields more useful models. We study a safety consequence of this mechanism in \emph{cross-session decomposition attacks}, where benign-looking subqueries are asked across independent interactions and later recomposed toward a forbidden objective. We formalize this setting as \emph{compositional safety risk} and prove a conditional risk-transfer bound: when the reference environment already contains dispersed evidence for a risky reconstruction, the gap between deployed composed risk and reference composed risk is controlled by the model's excess loss on allowed subqueries. Synthetic withholding experiments show that wider transformers assign lower loss to held-out instructions that never appear verbatim in training but are recoverable from injected supporting facts. A 600-intent pretrained-LLM evaluation shows that larger Qwen3 and Gemma3 family members can yield greater harmful-capability uplift under a fixed decomposition-composition pipeline. As a defense, IntentAlign-MiniLM, our 22M-parameter intent-aligned retriever, outperforms much larger embedding models on held-out intent retrieval and yields the best learned-retriever harmful recall across tested guardrails. Code is available in \href{https://github.com/liaodisen/Cross-Session-Decomposition-Attacks}{our GitHub repository}.
Comments30 pages, 3 figures, EMNLP 2026 Findings