面向深度代码模型的有效黑盒对抗攻击:基于结构与标识符扰动
Towards Effective Black-Box Adversarial Attacks on Deep Code Models via Structural and Identifier Perturbations
浏览论文内容
中文总结 AI 辅助
Strike框架通过输入条件的层次化扰动空间,结合LLM结构候选与标识符替换,在黑盒设置下有效攻击深度代码模型,提升攻击成功率并保持变体自然性。
中文摘要 AI 辅助
深度代码模型(DCM)日益嵌入代码智能任务中。然而,它们在对抗攻击下的鲁棒性仍未得到充分理解。先前的黑盒攻击主要依赖于仅标识符替换或从参考样本迁移的结构编辑,产生的扰动空间在很大程度上与被攻击输入无关。我们提出了Strike,一个输入条件的黑盒对抗鲁棒性测试框架,为每个输入构建并搜索一个层次化扰动空间。Strike首先将代码划分为块,并使用LLM生成上下文特定的结构候选,这些候选经过语法有效性过滤、按相似度排序并自适应组合。然后,它通过从动态构建的候选池中抽取的相似性引导的标识符替换来细化最佳结构变体。在代表性代码智能任务(包括克隆检测、漏洞检测和代码摘要)上的评估表明,Strike以具有竞争力的目标模型查询开销实现了更高的攻击成功率。针对所有三项任务的策略特定静态检查以及在可执行的Juliet漏洞检测子集上的基于执行的验证,为生成变体的有效性提供了互补的、任务受限的证据。表示相似性和人工评估进一步表明,这些变体与原始代码高度相似且在上下文中自然。使用Strike生成的样本进行微调,在固定的扰动评估集上改善了跨攻击性能,同时保持了干净性能。
英文摘要
Deep code models (DCMs) are increasingly embedded in code intelligence tasks. However, their robustness under adversarial attacks remains insufficiently understood. Prior black-box attacks mainly rely on identifier- only substitutions or structural edits transferred from reference samples, yielding perturbation spaces defined largely independently of the attacked input. We introduce Strike, an input-conditioned black-box adversarial robustness-testing framework that constructs and searches a hierarchical perturbation space for each input. Strike first partitions code into blocks and uses an LLM to generate context-specific structural candidates, filtered for syntactic validity, ranked by similarity, and adaptively combined. It then refines the best structural variant through similarity-guided identifier substitutions drawn from dynamically constructed candidate pools. Evaluations across representative code intelligence tasks, including clone detection, vulnerability detection, and code summarization, demonstrate that Strike achieves higher attack success rates with competitive target- model query overhead. Strategy-specific static checks across all three tasks and execution-based validation on the executable Juliet vulnerability-detection subset provide complementary, task-bounded evidence regarding the validity of the generated variants. Representation similarity and human evaluation further indicate that the variants remain highly similar to the original code and contextually natural. Fine-tuning with Strike- generated samples improves cross-attack performance on fixed perturbed evaluation sets while preserving clean performance.
发表机构
- The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。