TraceVIC:基于代码演化的因果推理识别漏洞引入提交
TraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits
浏览论文内容
中文总结 AI 辅助
TraceVIC通过时间图推理代码演化,识别并排序漏洞引入提交,在四个C/C++项目中将F2提升至0.814,比现有方法提高28.7%。
中文摘要 AI 辅助
软件漏洞通常在其被引入很久之后才被发现,这使得识别负责引入潜在漏洞条件的漏洞引入提交(VIC)变得困难。现有的VIC识别技术主要依赖git blame通过修订历史追踪漏洞代码,并使用位置启发式方法,例如选择其最早或最近的修改。然而,真正的VIC可能出现在该历史中的任何位置,并且漏洞行为可能依赖于跨多个修订演化的代码。因此,我们认为VIC识别需要对漏洞相关代码的演化方式进行推理,而不仅仅是候选提交在修订历史中出现的位置。我们提出了TraceVIC,一种基于时间图的方法,通过推理代码演化来识别和排序VIC。TraceVIC首先定位可能的根因代码行,并跨修订追踪其历史,构建捕获每个修订内程序结构以及漏洞相关代码在历史中演化的图表示。它利用时间边保留连续修订之间程序元素的对应关系,对得到的修订历史进行推理,并根据候选提交对漏洞条件的贡献直接对其进行排序。消融实验结果表明,建模完整修订历史将F2从0.637提升至0.814。TraceVIC在F2上比最先进的方法提高了高达28.7%,并在四个未见过的C/C++项目中,为79个漏洞中的78个识别出了有效的VIC。
英文摘要
Software vulnerabilities are often discovered long after they are introduced, making it difficult to identify the vulnerability-inducing commit (VIC) responsible for introducing the underlying vulnerable condition. Existing VIC identification techniques largely rely on git blame to trace vulnerable code through revision history and use positional heuristics, such as selecting its earliest or most recent modification. However, the true VIC may occur anywhere within this history, and vulnerable behavior may depend on code that evolves across multiple revisions. We therefore argue that VIC identification requires reasoning about how vulnerability-relevant code evolves, rather than simply where a candidate commit appears in the revision history. We present TraceVIC, a temporal graph-based approach for identifying and ranking VICs by reasoning over code evolution. TraceVIC first localizes likely root-cause lines and traces their histories across revisions, constructing graph representations that capture program structure within each revision and the evolution of vulnerability-relevant code across the history. It reasons over the resulting revision history, using temporal edges to preserve correspondences between program elements across consecutive revisions, and directly ranks candidate commits according to their contribution to the vulnerable condition. Ablation results show that modeling the full revision history improves F2 from 0.637 to 0.814. TraceVIC improves F2 by up to 28.7% over state-of-the-art methods and identifies a valid VIC for 78 of 79 vulnerabilities across four unseen C/C++ projects.
发表机构
- Northern Illinois University(北伊利诺伊大学)
- MIT Lincoln Laboratory(麻省理工学院林肯实验室)
机构由 AI 辅助整理,请以论文原文为准。