发表机构
University of Luxembourg(卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM编码智能体仅评估最终代码无法衡量安全问题演变的问题,提出安全债务线积分SDLI指标,结合多款SAST工具在多数据集上验证,为同步研究编码进展与安全风险提供新方法。
AI 中文摘要
LLM编码智能体在提交解决方案前会经历数百个中间代码状态。仅评估最终产物会导致安全问题的演变过程无法被衡量。我们提出了安全债务线积分(Security Debt Line Integral,SDLI),用于在智能体达到新的最佳测试通过率时累积静态分析风险。我们使用四款静态应用安全测试(SAST)工具实现了该指标,并对830次通过的SWE-bench运行产物、712个ProgramBench最终工作区以及13条公开的MirrorCode轨迹展开研究。其中两个大规模样本群体采用SDLI的终态特例。两款工具对通用缺陷枚举(Common Weakness Enumeration,CWE)类别的一致率在SWE-bench运行中为3.9%,在80次至少通过90%官方测试的ProgramBench运行中为26.2%。这些均为扫描器检测结果,并非经过验证的漏洞率。排除三个含大量咨询项的类别后,后者的比率降至6.2%。同一任务的不同运行测得的分数存在差异,而一次重构的ProgramBench运行暴露了其首次实现编写时就存在的持续性问题。一项修复案例研究在保留测试行为的同时降低了扫描器信号,但也揭示了其对等效API重写的敏感性。SDLI为同时研究进展与安全问题提供了一种方法,其在引导智能体和确认可利用漏洞方面的价值仍有待验证。
英文摘要
LLM coding agents can traverse hundreds of intermediate code states before submitting a solution. Evaluating only the final artifact leaves the evolution of security findings unmeasured. We introduce the Security Debt Line Integral (SDLI), which accumulates static-analysis risk when an agent reaches a new best test pass ratio. We instantiate it with four static application security testing (SAST) tools and study artifacts from 830 passing SWE-bench runs, 712 ProgramBench final workspaces, and 13 public MirrorCode trajectories. The two large populations use the final-state special case of SDLI. Two-tool Common Weakness Enumeration (CWE) class agreement occurs in 3.9% of SWE-bench runs and 26.2% of the 80 ProgramBench runs passing at least 90% of official tests. These are scanner findings, not validated vulnerability rates. Excluding three advisory-heavy classes reduces the latter rate to 6.2%. Same-task runs differ in their measured scores, while one reconstructed ProgramBench run exposes persistent findings from its first implementation write. A repair case study reduces the scanner signal while preserving tested behavior, but also reveals sensitivity to equivalent API rewrites. SDLI offers a way to study progress and security findings together. Its value for steering agents and confirming exploitable vulnerabilities remains to be established.