AI 中文总结
本文提出一种针对代码语言模型的可迁移对抗攻击,通过扰动代码标识符实现代码检索攻击,可使部分模型的MRR最高下降77%,凸显现有代码搜索方法的脆弱性。
AI 中文摘要
可靠的代码检索对开发者生产力和有效代码复用至关重要,但当前为搜索工具提供支持的神经代码语言模型(CLM)易受针对非功能文本元素的对抗攻击。本文提出一种与编程语言无关、可迁移的对抗攻击,利用CLM的这一漏洞,该方法在不改变代码片段功能的前提下,扰动其中的标识符,使代码人为地与目标查询对齐。研究表明,即便使用CodeT5+等较小的代码嵌入模型计算,该攻击也十分有效,且可迁移至Voyage-code-3等更大的闭源嵌入模型或Gemini-3.1-Pro等大语言模型;它能提升查询与任意无关代码片段的相似度,使最先进模型的关键检索指标如平均倒数排名(MRR)最高下降77%。实验结果凸显了当前代码搜索方法的脆弱性,强调需开发更稳健、语义感知的方法。
英文摘要
Reliable code retrieval is crucial for developer productivity and effective code reuse. However, current neural code language models (CLMs) powering search tools are susceptible to adversarial attacks targeting non-functional textual elements. In this paper, we introduce a programming language-agnostic, transferable, adversarial attack that exploits this CLM vulnerability. Our approach perturbs identifiers within a code snippet without altering the snippet's functionality to artificially align the code with a target query. We demonstrate that our attack, even when computed using smaller code embedding models, such as CodeT5+, is highly effective and transferable to larger, closed-source embedding models, like Voyage-code-3, or LLMs like Gemini-3.1-Pro. Our attack can increase the similarity between the query and arbitrary, irrelevant code snippets, consequently degrading key retrieval metrics such as the Mean Reciprocal Rank (MRR) of state-of-the-art models by up to 77%. The experimental results highlight the fragility of current code search methods and underscore the need for more robust, semantic-aware approaches.
CommentsEMNLP Findings 26