利用影响函数检测受污染的代码生成提示批次
Detecting Contaminated Code-Generation Prompt Batches via Influence Functions
- The University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出与威胁模型无关的CodeSIFT方法,利用影响函数检测受污染的代码生成提示批次,在3B至7B参数的三个代码LLM上,中高注入率下AUROC达0.98,性能优于静态分析基线。
AI中文摘要:
大型语言模型(LLM)越来越多地用于代码生成,但它们仍然容易受到引发不安全实现的提示的影响。现有的防御措施通常依赖于预定义的威胁模型或已知的漏洞模式,限制了它们对新型攻击的有效性。我们提出CodeSIFT,一种与威胁模型无关的检测方法,利用影响函数识别引发异常模型行为的提示批次。CodeSIFT不检测特定漏洞,而是测量生成代码的参数空间影响,并使用统计检验确定候选提示集是否偏离良性参考分布。为评估我们的方法,我们引入两个涵盖各种漏洞的基准数据集。我们在三个参数规模从3B到7B的开放权重代码LLM上评估CodeSIFT,在中高注入率下达到高达0.98的AUROC分数,同时保持校准良好的误报率,且显著优于静态分析基线。这些结果表明,基于影响函数的检测是识别恶意代码生成提示的有前景方向,无需底层攻击类别的先验知识。
英文摘要:
Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. Existing defenses typically rely on predefined threat models or known vulnerability patterns, limiting their effectiveness against novel attacks. We propose CodeSIFT, a threat-model-agnostic detection method that leverages influence functions to identify batches of prompts that induce anomalous model behavior. Rather than detecting specific vulnerabilities, CodeSIFT measures the parameter-space influence of generated code and uses a statistical test to determine whether a candidate prompt set deviates from a benign reference distribution. To evaluate our approach, we introduce two benchmark datasets covering a variety of vulnerabilities. We evaluate CodeSIFT on three open-weight code LLMs ranging from 3B to 7B parameters, achieving AUROC scores of up to 0.98 at moderate-to-high injection rates, while maintaining well-calibrated false positive rates and substantially outperforming static analysis baselines. These results suggest that influence-function-based detection is a promising direction for identifying malicious code-generation prompts without requiring prior knowledge of the underlying attack class.