arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoTrace:通过智能过程间探索从补丁到触发器

AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

Arastoo Zibaeirad, Marco Vieira, Thomas Zimmermann

arXiv 2607.12058首次发表:更新:

AI 中文总结

研究针对修复漏洞提交的触发器定位难题,提出AutoTrace智能管道逐层探索代码属性图定位漏洞触发器,并构建SinkTrace - Bench数据集。AutoTrace在基准测试中表现出色,还揭示了前沿语言模型在因果推理上的差距。

AI 中文摘要

给定一个修复漏洞的提交,触发器定位要找出将易受攻击的程序状态转变为具体不安全操作的特定语句。此问题比二进制漏洞检测更难,因为答案需要过程间的因果推理。我们提出了AutoTrace,一个智能管道,通过逐层探索代码属性图来定位漏洞触发器,由语言模型代理决定下一步查看位置,由确定性可接受门决定在报告触发器之前需要什么证据。代理不会自行接受触发器,每个报告的触发器都有来自图的明确证据支持。在完整的InterPVD基准测试中,AutoTrace达到了75.0%的VulnHit和80.8%的FuncHit,超过了同一语料库上的先前技术水平。基于相同机制,我们构建了SinkTrace - Bench数据集,该数据集将每个漏洞暴露为从攻击者控制的输入通过传播到危险操作的源到汇(S2S)因果链。在该数据集上对前沿语言模型进行基准测试,发现即使是最强的模型也难以区分匹配对,揭示了触发器定位所针对的因果推理差距。

英文摘要

Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the answer demands interprocedural, causal reasoning: in a substantial fraction of real-world CVEs the triggering statement lies several call layers outside the patched function, beyond the reach of static rule sets and pattern-matching language models alike. We present AutoTrace, an agentic pipeline that localizes vulnerability triggers by exploring a code property graph layer by layer, with LLM agents deciding where to look next and deterministic admissibility gates deciding what evidence is required before a trigger can be reported. Agents never accept a trigger on their own authority; every reported trigger is backed by explicit evidence drawn from the graph, so the pipeline covers both intra- and interprocedural vulnerabilities without relying on ungrounded model judgment. On the full InterPVD benchmark, AutoTrace reaches 75.0% VulnHit and 80.8% FuncHit, surpassing the prior state of the art on the same corpus. Building on the same machinery, we construct SinkTrace-Bench, a dataset that exposes each vulnerability as a source-to-sink (S2S) causal chain from attacker-controlled input through propagation to the dangerous operation, drawn from matched vulnerable and patched program states. It comprises 1,542 verifier-confirmed, perfectly balanced vulnerable/safe samples whose label fidelity we audit against expert annotations. Benchmarking frontier LLMs on it, we find that even the strongest struggle to separate the matched pairs, exposing the causal-reasoning gap that trigger localization targets. Artifact available at https://github.com/Erroristotle/AutoTrace.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑