发表机构
Faculty of Information Technology and Communication Sciences, Tampere University(坦佩雷大学信息与通信科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究实证分析了LLM智能体代码修复中的静默失败,发现遗漏、引入和不足三类共170例漏洞,揭示现有测试与审查机制不足,并提出改进评估方向。
AI 中文摘要
基于LLM的自动化代码修复智能体近年来在研究和软件工程实践领域均受到广泛关注。然而,对于通过语法和功能验证但仍保留或引入安全漏洞的补丁,关注却十分有限。本研究旨在系统性地识别和分类基于LLM的智能体代码修复中的此类静默失败。我们使用七个智能体框架配合GPT-4o-mini,在两个安全重点数据集SecurityEval和CVEfixes上进行了实证研究,共收集了1,030条有效执行轨迹。通过三轮定性编码和人工验证,确认了170例静默失败。主要结果如下:(i) 静默失败被分为三大类:遗漏(Omission)、引入(Introduction)和不足(Inadequacy)。遗漏占确认失败的48.2%,引入占30.6%,不足占21.2%。(ii) 在这三大类下细分出十种细粒度失败代码,展示了智能体在修复过程中如何遗漏必要的安全控制、应用不完整的防御措施或引入新的漏洞。(iii) 当前的测试通过评估和基于LLM的审查者角色在已确认案例中不足以暴露或拦截这些失败。(iv) 不同框架中出现了相似的不安全解决方案,表明可能存在共享的模型、提示或任务层面的影响,而单智能体和多智能体系统表现出不同的失败特征。本研究结果将帮助研究人员和从业者改进基于LLM的智能体代码修复评估,并开发超越功能正确性、覆盖所有生成产物的针对性验证方法。
英文摘要
LLM-based agents for automated code repair have received significant attention in recent years from both research and software engineering practice perspectives. However, limited attention has been paid to patches that pass syntactic and functional verification but still retain or introduce security vulnerabilities. The aim of this research is to systematically identify and categorize such silent failures in LLM-based agentic code repair. We conducted an empirical study using 1,030 valid execution traces produced by seven agent frameworks with GPT-4o-mini across two security-focused datasets, SecurityEval and CVEfixes. Through three iterations of qualitative coding and manual verification, 170 confirmed silent failures were identified. The key results are: (i) Three main categories of silent failures were identified: Omission, Introduction, and Inadequacy. Omission accounts for 48.2% of the confirmed failures, Introduction for 30.6%, and Inadequacy for 21.2%. (ii) Ten fine-grained failure codes were classified under these three categories, showing how agents omit required security controls, apply incomplete defenses, or introduce new vulnerabilities during repair. (iii) Current test-passing evaluation and LLM-based reviewer roles were insufficient to expose or intercept these failures in the confirmed cases. (iv) Similar insecure solutions appeared across different frameworks, suggesting possible shared model-, prompt-, or task-level influences, while single-agent and multi-agent systems showed different failure profiles. The results of this study will assist researchers and practitioners in improving the evaluation of LLM-based agentic code repair and developing targeted verification methods that go beyond functional correctness and cover all generated artifacts.