隐藏在评论中:代码大语言模型中的上下文注入攻击面
Hidden in the Comments: A Context-Injection Attack Surface in Code LLMs
浏览论文内容
中文总结 AI 辅助
本研究揭示代码大语言模型在推理时易受上下文注入攻击,通过嵌入恶意注释可诱导其生成含漏洞代码,攻击成功率高达77%以上,且指令微调与模型规模影响有限,凸显了来源感知训练的必要性。
中文摘要 AI 辅助
代码大语言模型(Code LLM)助手从异构的开发上下文中生成代码,包括打开的文件、导入的模块、粘贴的代码片段和注释,其中许多内容可能来自不受信任的来源。我们研究了嵌入在这种上下文中的不安全指令是否能够在无需访问模型权重或训练数据的情况下,引导代码大语言模型生成易受攻击的代码。我们评估了十个开放权重的代码大语言模型,参数规模从3B到13B不等,包括四个基础模型和六个指令微调模型,覆盖十个Web应用弱点类别。我们将包含以代码注释形式嵌入的不安全指令的补全任务与不包含恶意指令的良性任务进行比较。在攻击条件下,补全结果中包含中等或更高弱点的情况占77.4%至92.3%,而在良性条件下这一比例为1.7%至5.1%。基础模型和指令微调模型平均分别产生86.5%和84.5%的易受攻击输出;等价性检验和三个匹配的模型对表明,指令微调后最多减少8.1%。易感性(susceptibility)与模型规模或专业化程度没有明显关联。在易受攻击的攻击输出中,86.2%至91.0%被评为高或严重等级,并且在没有基于模式的检测器的情况下,这种效应仍然存在。生成后的筛选降低了风险但未完全消除风险,最强的筛选仍留下约三分之一的未检测到的漏洞。这些发现将推理时的上下文注入确定为一种重大的攻击面,并激励了具有来源感知的训练目标。
英文摘要
Code large language model (Code LLM) assistants generate code from heterogeneous development contexts, including open files, imported modules, pasted snippets, and comments, much of which may originate from untrusted sources. We investigate whether insecure instructions embedded in such contexts can steer Code LLMs toward vulnerable code without access to model weights or training data. We evaluate ten open-weight Code LLMs spanning 3B--13B parameters, including four base and six instruction-tuned models, across ten web-application weakness classes. We compare completion tasks containing insecure instructions embedded as code comments with benign tasks without malicious instructions. Attack-condition completions contained a medium-or-higher weakness in {\bf 77.4--92.3}\% of cases, compared with {\bf 1.7--5.1}\% in the benign condition. Base and instruction-tuned models averaged 86.5\% and 84.5\% vulnerable outputs, respectively; equivalence testing and three matched model pairs indicated reductions of at most 8.1\% after instruction tuning. Susceptibility showed no clear association with model scale or specialization. Among vulnerable attack outputs, 86.2--91.0\% were rated high or critical, and the effect persisted without the pattern-based detector. Post-generation screening reduced but did not eliminate the risk, the strongest screen leaving roughly one-third undetected. These findings identify inference-time context injection as a substantial attack surface and motivate provenance-aware training objectives.
发表机构
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。