发表机构
University of Duisburg-Essen(杜伊斯堡-埃森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CodePoisonRAG框架,可将良性代码构件转化为投毒构件,对检索增强代码生成(RACG)发起有针对性的上游知识投毒攻击,在3种生成器及防御工具CodeGuarder上均取得较高攻击成功率。
AI 中文摘要
检索增强代码生成(Retrieval-Augmented Code Generation,RACG)通过检索外部代码构件、文档和补丁并将其融入生成上下文,改进了基于大语言模型(LLM)的软件开发。这种对外部知识的依赖引入了关键的信任边界:被投毒的构件可在不修改底层LLM的情况下影响生成的代码。现有研究表明,选择存在漏洞的现有示例可提高RACG输出的整体漏洞率,但未解决黑盒攻击者能否构建单个任务匹配的构件以传播攻击者选定的弱点这一问题。我们提出CodePoisonRAG,这是一种有针对性的上游知识投毒框架,可将良性固定代码条目转化为被投毒的构件。其攻击链结合了特定通用弱点枚举(Common Weakness Enumeration,CWE)的漏洞注入(嵌入选定的源到汇流,同时保持任务一致性)与语义错误标注(添加虚假安全声明但不修复漏洞行为)。攻击者无法访问受害者部署的知识库、检索器、重排序器、生成器、提示或防御机制,且针对每个预期编程任务最多注入1个构件。我们构建了覆盖Java和C语言10个CWE类别的85个被投毒构件,总语料投毒率为0.7%。在3种生成器上,所有85个构件均出现在对应查询的前3个结果中,CodePoisonRAG的攻击成功率介于0.80至0.93之间。针对在生成上下文注入漏洞特定安全知识的CodeGuarder,攻击仍保持0.40至0.71之间的成功率。这些结果表明,RACG投毒不仅限于现有漏洞的偶然传播,还可实现攻击者选定弱点的针对性构建与传播。
英文摘要
Retrieval-Augmented Code Generation (RACG) improves LLM-based software development by retrieving external code artifacts, documentation, and patches, and incorporating them into the generation context. This reliance on external knowledge introduces a critical trust boundary: poisoned artifacts can influence generated code without modifying the underlying LLM. Prior work shows that selecting existing vulnerable examples can increase the general vulnerability rate of RACG outputs, but leaves open whether a black-box attacker can construct a single task-matched artifact that propagates an attacker-selected weakness. We introduce CodePoisonRAG, a targeted upstream knowledge-poisoning framework that transforms benign fixed-code entries into poisoned artifacts. Its attack chain combines CWE-specific Vulnerability Injection, which embeds a selected source-to-sink flow while retaining task alignment, with Semantic Mislabeling, which adds false safety claims without repairing the vulnerable behavior. The attacker has no access to the victim's deployed knowledge base, retriever, re-ranker, generator, prompt, or defense mechanism and injects at most one artifact per anticipated programming task. We construct 85 poisoned artifacts covering ten CWE classes across Java and C, yielding an aggregate corpus-poisoning ratio of 0.7%. Across three generators, all 85 artifacts appear among the Top-3 results for their corresponding queries, and CodePoisonRAG achieves attack success rates between 0.80 and 0.93. Against CodeGuarder, which injects vulnerability-specific security knowledge into the generation context, the attack retains success rates between 0.40 and 0.71. These results show that RACG poisoning extends beyond the incidental propagation of existing vulnerabilities to the targeted construction and propagation of attacker-selected weaknesses.
Comments16 pages, 1 figure. Under review