arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探测预填充:通过潜在激活检测代码漏洞

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

Alizishaan Khatri

arXiv 2608.16970首次发表:更新:

发表机构

Wrynx Inc.(瑞恩克斯公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从四个LLM提取预填充标记激活训练MLP探测器,在四个代码漏洞基准上实现平均F1值41.7%,证明LLM自身代码表示含漏洞信息,为轻量原生筛查提供方向。

AI 中文摘要

基于大语言模型(LLM)的代码生成现已嵌入关键任务流程,但针对易受攻击输出的防御措施仍为事后应对——静态分析器、微调分类器或LLM评估器仅在代码生成完成后进行筛查,忽略了生成模型自身的内部状态。我们检验一个更狭窄、可直接测量的问题:当LLM将一段C/C++代码作为上下文读取时,其隐藏激活是否已携带该代码漏洞状态的信号?我们从三个模型家族的四个LLM(Granite-4.1-8B、Qwen3.5-9B、Qwen3.6-27B、Gemma-4-12B)中提取最后预填充标记的激活,并在这些激活上训练多层感知机(MLP)探测器。我们在四个函数级C/C++基准数据集(Devign、Big-Vul、Draper VDISC、PrimeVul)上对探测器进行评估。我们的探测器使用参数规模为1340万至1600万的探测器,平均F1值达41.7%,仅为基础模型规模的0.2%以下。在Devign数据集上,表现最佳的探测器(Qwen3.5-9B,F1值68.8%)与已发表的微调分类器的最优结果(SOTA,67.9%)相当,尽管仅读取了一个冻结的通用LLM的激活;在更困难、类别更不平衡的基准数据集(Big-Vul、Draper VDISC、PrimeVul)上,探测器的表现则大幅落后于最优结果。这是早期证据,表明编码LLM对任意代码的自身表示包含该代码漏洞状态的信息,为开发轻量级、模型原生的漏洞筛查方法提供了研究动力。

英文摘要

LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. We test a narrower, directly measurable question: when an LLM reads a piece of C/C++ code as context, do its hidden activations already carry a signal about that code's vulnerability status? We extract last prefill token activations from four LLMs (Granite-4.1-8B, Qwen3.5-9B, Qwen3.6-27B, Gemma-4-12B) across three model families and train MLP probes on these activations. We evaluate them on four function-level C/C++ benchmarks (Devign, Big-Vul, Draper VDISC, PrimeVul). Our probes achieve 41.7\% average F1 using 13.4--16.0M-parameter probes -- under 0.2\% of base-model size. On Devign, the best probe (Qwen3.5-9B, 68.8\% F1) matches the published fine-tuned-classifier SOTA (67.9\%) despite reading only a frozen, general-purpose LLM's activations; on the harder, more imbalanced benchmarks (Big-Vul, Draper VDISC, PrimeVul) probes trail SOTA substantially. This is early evidence that a coding LLM's own representation of arbitrary code is informative about that code's vulnerability status, motivating further work toward lightweight, model-native vulnerability screening.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑