arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IssueTrojanBench:针对恶意问题请求对人工智能编码代理进行基准测试

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

Ankur Singh, Jinqiu Yang, Tse-Hsun Chen

arXiv 2607.20759首次发表:更新:

发表机构

Concordia University(康科德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对恶意问题请求对人工智能编码代理进行基准测试,通过构建包含多种攻击类别和交付向量的新型基准IssueTrojanBench,发现现代编码代理存在关键漏洞,强调需更强安全机制保护。

AI 中文摘要

由大语言模型驱动的人工智能编码代理越来越多地集成到实际软件开发中,它们可以自主访问本地文件和工具来生成、编辑和执行代码。编码代理继承了来自大语言模型主干和其代理架构的安全风险。本文对由两个主要模型家族驱动的最先进编码代理(Cursor、Claude Code和Codex Desktop)进行了针对恶意问题请求的系统评估。我们的新型基准测试IssueTrojanBench包含基于四种新型攻击类别构建的恶意问题,通过六种交付向量,并通过扰动进一步增强。结果揭示了现代编码代理中的关键漏洞,即来自IssueTrojanBench的66.5%的恶意问题穿透了编码代理的所有防护栏(代理和大语言模型级别)。进一步分析表明,拒绝几乎完全来自大语言模型而非代理框架,GPT模型普遍易受攻击,Sonnet 4.6对高影响行动表现出更具选择性、有风险意识的阻止。我们的评估还强调当前代理级防御策略为编码代理提供的额外保护有限。我们的发现突出了迫切需要更强的代理和模型级安全机制来保护人工智能编码代理。

英文摘要

AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with autonomous access to local files and tools. Coding agents inherit security risks from both the LLM backbone, where adversarial prompts, poisoned training data, and backdoor triggers can cause models to emit insecure or attacker-chosen code, and their agentic architecture, where tool-using autonomy enables induced misuse of external APIs, data exfiltration, and persistent compromise of development environments. This paper presents a systematic evaluation of malicious issue requests against state-of-the-art coding agents (Cursor, Claude Code, and Codex Desktop), powered by two major model families (OpenAI GPT-5.3 Codex/GPT-5.4 and Anthropic Sonnet 4.6). Our novel benchmark IssueTrojanBench contains malicious issues that are constructed based on four novel attack categories (i.e., embedded as malicious instructions in issues), six delivery vectors (e.g., PDF, or issue comment), and further augmented by perturbations. Our results reveal critical vulnerabilities in the as-deployed modern coding agents, i.e., 66.5% of the malicious issues from IssueTrojanBench penetrate all the guardrails (agent- and LLM-level) of coding agents. Our further analysis shows that rejection is almost entirely from LLMs rather than the agent frameworks, with GPT models broadly vulnerable and Sonnet 4.6 exhibiting more selective, risk-aware blocking of high-impact actions. Our evaluation also highlights that the current agent-level defense strategy offers limited additional protection for coding agents. Our findings highlight the urgent need for stronger agent- and model-level safety mechanisms to protect AI coding agents.

Comments10 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑