arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

激活探针:挖掘模型输出遗漏的表面代码安全信号

Activation Probes Surface Code-Security Signals that the Model's Output Misses

Ivan Wiryadi

arXiv 2608.09643首次发表:更新:

AI 中文总结

本研究提出用线性探针分析开源审查器模型的激活值,发现其携带了仅通过模型提示无法获取的代码安全信号,可有效区分有漏洞与修复后的函数。

AI 中文摘要

当前AI编码代理生成的生产代码占比不断提升,人工安全审查的速度无法匹配代码生成速度。主流AI编码代理为闭源模型,部署团队无法查看其内部结构,因此可使用开源模型作为审查器,对代理的输出进行审查,而该审查器的激活值是可读的。本文研究读取这些激活值是否能恢复出仅通过向同一审查器提问所遗漏的安全信号。研究人员在包含有漏洞及修复后Python函数的配对语料库上,为每个模型拟合了单个线性探针,随后在5种开源审查器模型上,对探针训练时从未见过的漏洞类型的真实公开漏洞进行测试,且测试过程中不重新训练探针。在通过修改单个函数修复的漏洞中,所有模型的探针都能在61%-67%的案例中,将有漏洞函数的评分高于其修复后的函数,优于50%的随机概率;且在所有测试的提示下,探针的表现也优于从模型logits中读取的同一模型的提示式是/否胜率。向模型询问书面判断,即便采用思维链(chain-of-thought),多数情况下对有漏洞函数和修复后函数返回相同答案,无法区分二者。研究表明,模型激活值携带了仅通过提示同一模型所遗漏的代码安全信号。

英文摘要

AI coding agents now write a growing share of production code, and human security review does not scale at the rate code is generated. The agents in widest use are closed-weight, so a deploying team cannot read their internals. It can instead run an open-weight model as a reviewer over the agent's output. That reviewer's activations are readable. We ask whether reading those activations recovers a security signal that simply asking the same reviewer misses. We fit a single linear probe per model on a corpus of paired vulnerable-and-fixed Python functions, then test it without retraining on real disclosed vulnerabilities whose weakness type the probe never saw in training, across five open-weight reviewer models. On the vulnerabilities fixed by changing a single function, the probe scores the vulnerable function above its fix on 61-67% of cases for every model, beating the 50% chance line. It also beats the same model's prompted YES/NO win-rate read from its logits, under every prompt we try. Asking the model for a written verdict, even with chain-of-thought, returns the same answer on the vulnerable and fixed function most of the time and so cannot tell them apart. Model activations carry a code-security signal that prompting the same model misses.

Comments6 pages, 1 figure. Accepted at the TAIGR workshop, ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑