发表机构
Qingdao University of Technology; Ocean University of China(青岛理工大学; 中国海洋大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GRAML框架通过图证据引导GPT-5进行思维树推理并生成漏洞描述,结合多任务训练,在漏洞检测上平均F1达66.67%-68.70%,优于现有基线最多30.92%。
AI 中文摘要
大型语言模型(LLMs)已被广泛应用于软件漏洞检测。然而,其性能往往因对控制流和数据流信息的利用不足而受到限制。在本文中,我们提出了GRAML,一个结合图证据、漏洞描述生成和多任务训练的框架。GRAML首先对C/C++程序进行静态分析,提取关键源代码行和带类型的行间关系作为结构证据。然后,它利用这些证据引导GPT-5通过思维树引导的漏洞推理(ToT-VR)过程,并生成漏洞描述。这些描述进一步与检测、定位和评估样本相结合,构建一个统一的四任务训练数据集。我们在分布内(ID)测试集和六个分布外(OOD)数据集上评估了GRAML。结果表明,GRAML的平均F1分数在66.67%到68.70%之间,比最先进的基线高出最多30.92%。消融实验进一步表明,与标准思维链(CoT)推理和原始代码属性图(CPG)序列化相比,ToT-VR和图引导的漏洞描述提高了检测性能。这些发现为使用大型语言模型构建更可靠、更安全的软件工程系统提供了实践指导。
英文摘要
Large Language Models (LLMs) have been widely applied to software vulnerability detection. However, their performance is often limited by insufficient use of control-flow and data-flow information. In this paper, we propose GRAML, a framework that combines graph evidence, vulnerability description generation, and multi-task training. GRAML first performs static analysis on C/C++ programs to extract critical source lines and typed line relations as structural evidence. It then uses this evidence to guide GPT-5 through the Tree-of-Thought-guided Vulnerability Reasoning (ToT-VR) process and generate vulnerability descriptions. These descriptions are further combined with Detection, Localization, and Assessment samples to build a unified four-task training dataset. We evaluate GRAML on an in-distribution (ID) test set and six out-of-distribution (OOD) datasets. The results show that GRAML achieves average F1 scores ranging from 66.67% to 68.70%, outperforming state-of-the-art baselines by up to 30.92%. Ablation experiments further show that ToT-VR and graph-guided vulnerability descriptions improve detection performance compared with standard Chain-of-Thought (CoT) reasoning and raw Code Property Graph (CPG) serializations. These findings provide practical guidance for building more reliable and secure software engineering systems with large language models.