AI 中文总结
研究如何通过对抗性代码注释对基于大语言模型的漏洞检测器进行自适应智能体攻击,提出ALIBI框架,将现实世界漏洞修复提交转换为编码任务评估多个检测器,发现其高度易受攻击,揭示攻击面并推动安全意识设计。
AI 中文摘要
大语言模型越来越多地用于漏洞检测和代码审查等安全敏感任务。它们对嵌入在源代码中的自然语言上下文的依赖暴露了一个以前未被充分探索的攻击面:即可以在不改变程序行为的情况下影响检测器推理的对抗性注释。我们研究针对一种新对手的基于大语言模型的漏洞检测器:一种实现新功能、故意引入漏洞并策略性插入对抗性源代码注释以逃避检测的编码智能体。我们提出了ALIBI,一个自动化的自适应黑盒攻击框架,它利用检测器的推理和反馈来生成并迭代优化对抗性注释。我们将现实世界中的漏洞修复提交转换为编码任务,并评估了四个具有代表性的基于大语言模型的漏洞检测器,从专门的开放权重推理模型到前沿的多智能体系统。所有评估的检测器都高度易受攻击:在125个现实世界的空指针解引用漏洞中,攻击成功率超过90%,在一个系统上达到100%。该框架也能推广到其他漏洞类别。引导检测器推理或伪造外部工具结果的对抗性注释最为有效,而基于检测器反馈的迭代优化进一步提高了攻击成功率。最后,提示级防御对自适应攻击的鲁棒性有限,而架构隔离和预检测器注释清理则显著提高了弹性。我们的发现揭示了当前基于大语言模型的漏洞检测器中一个基本的攻击面,并推动了安全意识设计,仔细校准自然语言上下文和程序证据之间的信任。
英文摘要
Large language models are increasingly deployed for security-sensitive tasks such as vulnerability detection and code review. Their reliance on natural-language context embedded in source code exposes a previously underexplored attack surface: adversarial comments that can influence a detector's reasoning without changing program behavior. We study LLM-based vulnerability detectors against a new adversary: a coding agent that implements new functionality, deliberately introduces vulnerabilities, and strategically inserts adversarial source-code comments to evade detection. We present ALIBI, an automated adaptive black-box attack framework that generates and iteratively refines adversarial comments using detector reasoning and feedback. We transform real-world vulnerability-fixing commits into coding tasks and evaluate four representative LLM-based vulnerability detectors, ranging from specialized open-weight reasoning models to frontier multi-agent systems. All evaluated detectors are highly vulnerable: attack success rates exceed 90% across 125 real-world null-pointer dereference vulnerabilities, reaching 100% on one system. The framework also generalizes beyond this vulnerability class. Adversarial comments steering detector reasoning or fabricating external tool results prove most effective, while iterative refinement based on detector feedback further increases attack success. Finally, prompt-level defenses provide limited robustness against adaptive attacks, whereas architectural isolation and pre-detector comment sanitization substantially improve resilience. Our findings expose a fundamental attack surface in current LLM-based vulnerability detectors and motivate security-aware designs that carefully calibrate trust between natural-language context and program evidence.