arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动化程序修复智能体针对安全漏洞的对抗性测试

Adversarial Testing of Automated Program Repair Agents for Security Vulnerabilities

Fares Trad, Simin Chen, Hung Viet Pham, Gias Uddin, Baishakhi Ray

arXiv 2609.15963首次发表:更新:

AI 中文总结

本研究通过构建SWEADV基准并测试多个APR智能体,发现对抗性问题描述可诱导51.7%的恶意修复,现有检测机制准确率低,表明自主APR智能体尚不能安全用于生产。

AI 中文摘要

配备大型语言模型(LLM)的软件智能体被设计用于自动化程序修复(APR)任务,这引发了在不久的将来APR智能体将在无需太多人工干预的情况下自动修复错误的可能性。在这种情况下,我们能否信任APR智能体生成既功能正确又安全的代码?如果攻击者用看似良性但可能影响智能体生成正确但不安全代码的对抗性问题来针对生产环境中的APR智能体,又会怎样?在本文中,我们通过进行一项实证研究,迈出了回答这些问题的第一步。首先,我们创建了SWEADV,这是一个由SWE-bench Verified中150个修复任务构建的包含750个对抗性问题描述的基准。对于每个修复任务,我们创建了五个对抗性问题描述,每种攻击类型一个:命令执行、反序列化、路径遍历、拒绝服务和弱哈希。其次,我们在SWEADV上评估了来自三个LLM后端的mini_swe APR智能体:GPT-5-Mini、MiniMax-M2.5和DeepSeek-R。我们发现,平均而言,对抗性问题描述可以诱导恶意行为,并在51.7%的情况下成功修复。第三,我们调查了典型的检测机制是否足以阻止此类恶意补丁被接受。使用LLM作为评判者对对抗性问题描述进行修复前检测,平均检测准确率仅为62.3%。使用静态分析工具和LLM作为评判者对对抗性APR补丁进行修复后检测,平均检测准确率分别仅为39.4%和55.4%。我们得出结论,鉴于自主APR智能体易受对抗性攻击,在生产部署中尚不能信任它们。

英文摘要

Software agents with Large Language Models (LLMs) are designed for Automated Program Repair (APR) tasks, raising the possibility that, in the near future, APR agents will fix bugs automatically without much human intervention. Can we trust an APR agent to produce both functionally correct and secure code in such situations? What if attackers target production APR agents with adversarial issues that seem benign but may influence the agents to produce correct but insecure code? In this paper, we took a first step towards answering these questions by conducting an empirical study. First, we created SWEADV, a benchmark of 750 adversarial issue descriptions constructed from 150 repair tasks in SWE-bench Verified. For each repair task, we created five adversarial issue descriptions, one for each attack type: command execution, deserialization, path traversal, denial of service, and weak hashing. Second, we evaluated mini_swe APR agents from three LLM backends on SWEADV: GPT-5-Mini, MiniMax-M2.5, and DeepSeek-R. We found that on average, adversarial issue descriptions can induce malicious behaviors with successful repair in 51.7% of cases. Third, we investigated whether typical detection mechanisms are sufficient to prevent such malicious patches from being accepted. Pre-repair detection with LLM-as-judge on the adversarial issue descriptions resulted in an average detection accuracy of only 62.3%. Post-repair detection on adversarial APR patches using static analysis tools and LLM-as-judge achieved average detection accuracies of only 39.4% and 55.4%, respectively. We conclude that autonomous APR agents cannot be trusted yet in production deployment, given their susceptibility to adversarial attacks.

Comments12 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑