arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29775cs.CRcs.AI

预填充推理通道:针对推理型大语言模型的输出前缀攻击

Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs

Lukáš Brůna, Robert Bridges, Adam Ek

首次发表
浏览论文内容

中文总结 AI 辅助

本研究首次系统隔离推理通道作为输出前缀攻击向量,发现仅注入恶意推理无效,但结合输出前缀可使攻击成功率高达99%,且效果因模型而异。

中文摘要 AI 辅助

大型语言模型(LLMs)消费并产生单一文本序列;因此,如果可以在LLM响应的开头添加文本,即输出前缀,那么所有后续标记都将以此为前提条件。这种输出前缀攻击技术是一种廉价的黑盒提示注入。先前工作已表明,此类攻击能可靠地越狱非推理模型。大多数推理模型在助手的最终响应之前增加了一个中间草稿本推理步骤。一些API暴露了编辑此推理通道的能力,攻击向量可用于推理注入攻击。我们首次进行了系统的、受控的研究,将草稿本推理通道隔离为输出前缀攻击向量,并首次在暴露推理和隐藏推理模型上比较了仅推理、仅输出前缀以及推理加输出前缀的攻击。采用3种前缀类型×2种推理注入的因子设计,基于AdvBench的1800多个测试用例,我们攻击了三个2026时代的前沿模型:Gemini 3 Flash Preview、DeepSeek V4 Flash和Claude Haiku 4.5。我们发现,单独注入恶意推理基本上无效(攻击成功率约为0%),但将相同推理与一个平凡的输出前缀一起注入,可将某些模型的攻击成功率提高到高达99%。对于此类攻击,我们发现上下文前缀比静态前缀更有效;且易感性取决于模型。

英文摘要

Large Language Models (LLMs) consume and produce a single sequence of text; hence, if text can be added to the beginning of the LLM's response, i.e., an output prefix, then all subsequent tokens will be conditioned on it. This output-prefix attack technique is a cheap black-box prompt injection. Prior work has shown this type of attack can reliably jailbreak non-reasoning models. Most reasoning models add an intermediate scratchpad reasoning step before the assistant's final response. The ability to edit this reasoning channel is exposed by some APIs and attack vectors can be leveraged for reasoning injection attacks. We present the first systematic, controlled study that isolates the scratchpad reasoning channel as an output-prefix attack vector, and the first to compare reasoning-only, output-prefix-only and reasoning-plus-output-prefix attacks across both exposed- and hidden-reasoning models. Using a factorial design of 3 prefix types $\times$ 2 reasoning injections over $1{,}800$ test cases drawn from AdvBench, we attack three 2026-era frontier models Gemini 3 Flash Preview, DeepSeek V4 Flash, and Claude Haiku 4.5. We find that injecting malicious reasoning alone is essentially inert ($\approx0\%$ attack success), but injecting the same reasoning together with a trivial output prefix raises the attack success rate to as high as $99\%$ for some models. For this type of attack we find that contextual prefixes work better than static prefixes; and that susceptibility is dependent on the model.

发表机构

  • Uppsala University(乌普萨拉大学)
  • AI Sweden(瑞典人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑