发表机构
Capital One(第一资本金融公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究将工作空间拓扑作为攻击向量,发现其会影响智能体代码助手的间接提示注入攻击成功率,高度模块化环境及安全提示可降低攻击成功率,为代码智能体安全测试提供实用价值。
AI 中文摘要
智能体代码助手不仅用于开发新代码,还广泛用于快速摄取和利用第三方代码,这带来了恶意代码被摄取的风险,因为这些代码工具在开发者工作空间内拥有广泛的文件系统访问权限。本文对一种新型攻击面的不同维度的影响进行了广泛研究,该攻击面被称为工作空间拓扑,其定义包括目录深度、代码库模块化程度、文件内注入位置和上下文框架,研究其对对抗性提示注入尝试攻击成功率的影响。我们对间接提示注入(IPI)进行了实证研究,涉及10种编程语言和6个工程领域的多样化开源代码库,针对使用开源代码工具的开放权重模型评估了三个IPI入口点。我们发现工作空间拓扑会显著影响IPI成功率,具体而言,代码库模块化程度的变化可显著改变攻击成功率(ASR),高度模块化环境的攻击成功率明显更低;此外,工作空间中的上下文框架和安全提示的引入也会改变ASR。我们的研究结果为不同场景下代码智能体的评估和安全测试提供了实用价值,同时强调了使用未受污染的测试环境以获得可靠结果和结论的重要性。
英文摘要
Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party code. This opens up a risk of malicious code being ingested as these coding tools operate with broad filesystem access inside developer workspaces. In this paper, we extensively study the impact of different dimensions of a novel attack surface we term workspace topology -- defined via directory depth, codebase modularity, in-file injection position and context framing -- on the attack success rate of adversarial prompt injection attempts. We perform an empirical study of indirect prompt injection (IPI) across a diverse set of open-source repositories spanning 10 languages and 6 engineering domains, evaluating three IPI entry points against open-weight models operating open source code harnesses. We find that workspace topology measurably affects IPI success. Specifically, changes in codebase modularity can significantly alter the Attack Success Rate (ASR), with highly modular environments demonstrating significantly lower attack success rates. Furthermore, context framing and introduction of security-cues in the workspace can also alter the ASR. Our findings offer practical value for the evaluation and security testing of coding agents across diverse settings, while underscoring the importance of an uncontaminated testing environment to obtain reliable results and conclusions.
Comments15 pages, 10 figures. Preprint of a paper accepted at the Conference on Applied Machine Learning in Information Security (CAMLIS 2026)