arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

修复最初在哪里出错?定位智能体漏洞修复中静默失败的起源

Where Did the Repair First Go Wrong? Localizing the Origins of Silent Failures in Agentic Vulnerability Repair

Wenji Bai, Muhammad Waseem, Zeeshan Rasheed, Jaakko Peltonen, Pekka Abrahamsson

arXiv 2610.06163首次发表:更新:

发表机构

Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体漏洞修复中的静默失败,提出基于轨迹的SAGE方法,通过分析每轮安全推理与代码历史,定位最早偏离安全意图的环节,在95个案例中成功定位93个,多数源于未解决的安全要求或不当防御。

AI 中文摘要

定位基于大语言模型的智能体在修复过程中首次未能维护安全性的位置,可以揭示其工作流程的哪个阶段需要额外的防护措施。这对于静默失败而言尤为困难,静默失败是指通过语法和功能检查但仍包含安全漏洞的补丁。由于此类补丁不产生可观察的失败信号,现有的失败归因方法依赖于观察到的任务失败和标记的失败步骤,因此不太适用于它们。我们提出安全感知差距评估(SAGE),一种基于轨迹的方法,结合对每一轮记录的安全推理的评估与重建的代码历史,以识别修复偏离任务安全意图的最早轮次。我们在从六个智能体框架和六个基础模型在SecurityEval和CVEfixes上产生的3,684条修复轨迹中抽取的95个已确认的静默失败上评估SAGE。SAGE在93个案例中分配了起源。大多数起源是未解决的安全要求或不充分的防御选择,只有五个与代码更改本身重合。当智能体引入易受攻击的代码时,在19个案例中有14个案例中起源先于写入。重复评分和第二位评判员在再现起源类型方面比精确轮次更一致,而仅保留最终文件的轨迹的一致性最低。

英文摘要

Localizing where an LLM-based agent first fails to uphold security during a repair can show which stage of its workflow needs an additional safeguard. This is difficult for silent failures, which are patches that pass syntactic and functional checks but still contain a security vulnerability. Because such patches give no observable failure signal, existing failure attribution methods, which rely on observed task failures and labelled failure steps, are less suited to them. We propose Security Awareness Gap Evaluation (SAGE), a trace-based method that combines an assessment of the security reasoning recorded at each turn with the reconstructed code history to identify the earliest turn at which a repair diverges from the task's security intent. We evaluate SAGE on 95 confirmed silent failures drawn from 3,684 repair traces produced by six agent frameworks and six base models on SecurityEval and CVEfixes. SAGE assigned an origin in 93 cases. Most origins were an unaddressed security requirement or an inadequate defence choice, and only five coincided with the code change itself. When the agent introduced the vulnerable code, the origin preceded the write in 14 of 19 cases. Repeated scoring and a second judge reproduced the origin type more consistently than the exact turn, and agreement was lowest for traces that kept only the final file.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑