AI 中文总结
本文利用AIDev数据集,通过大语言模型评判框架和人工分析,研究自主编码代理生成拉取请求中的安全代码异味。发现38.9%的请求含安全异味,供应链问题和硬编码凭证占比高,揭示了安全风险及开发者警惕性降低问题,强调需实施上下文感知安全护栏。
AI 中文摘要
自主编码代理的日益采用加速了软件开发,但也在高影响文件路径中引入了超出传统人工审查能力的安全风险。此前研究主要评估这些系统的功能正确性和生产力,本文利用AIDev数据集进行大规模实证研究,以系统地描述代理生成的拉取请求中的安全代码异味。通过经过验证的基于大语言模型的评判框架和人工定性分析相结合的方式,我们识别并分类了跨越4022个拉取请求的16112个文件更改中的安全配置错误。结果显示,38.9%的代理生成的拉取请求至少包含一种安全异味,供应链完整性问题占所有检测到的安全异味的82.3%。此外,硬编码凭证占所有严重安全异味的99.6%。关键的是,我们发现人类合作者在这些代理辅助工作流程中导致了67.6%的真正秘密泄露,而现有的自动化和人工审查流程在集成前未能检测到81.1%的这些凭证。这些发现突出了代理辅助软件开发工作流程中的重大安全风险,并表明开发者警惕性可能降低。它们还强调了在人机协作点直接实施上下文感知安全护栏的迫切需求。
英文摘要
The increasing adoption of autonomous coding agents accelerates software development but also introduces scoped security risks within high-impact file paths that can outpace traditional human review capacity. While prior research has primarily evaluated these systems in terms of functional correctness and productivity, this paper presents a large-scale empirical study using the AIDev dataset to systematically characterize security code smells in agent-generated pull requests (PRs). Through a combination of a validated LLM-as-a-judge framework and manual qualitative analysis, we identify and classify security misconfigurations across 16,112 file changes spanning 4,022 pull requests. Our results reveal that 38.9% of agent-generated PRs contain at least one security smell, with supply chain integrity issues accounting for 82.3% of all detected security smells. Furthermore, hard-coded credentials constitute 99.6% of all critical-severity security smells. Crucially, we find that human collaborators are responsible for introducing 67.6% of genuine leaked secrets within these agent-assisted workflows, while existing automated and human review processes fail to detect 81.1% of these credentials prior to integration. These findings highlight substantial security risks in agent-assisted software development workflows and suggest a potential reduction in developer vigilance. They also underscore the urgent need for context-aware security guardrails implemented directly at the point of human-AI collaboration.
CommentsAccepted at the KDD 2026 Workshop on Agentic Software Engineering (AgenticSE)