使用混合网络模型桥接软件漏洞检测的语义与结构
Bridging Semantics & Structure for Software Vulnerability Detection using Hybrid Network Models
- The George Washington University(乔治·华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出将程序异构图表示与轻量级本地LLM结合的混合框架,在Java漏洞检测中达93.57%准确率,同时提取子图生成解释,提升可解释性。
AI中文摘要:
软件漏洞仍是持续存在的风险,但静态和动态分析经常忽略塑造不安全行为的结构依赖。我们将程序视为异构图,捕获控制流和数据流关系作为复杂交互网络。我们的混合框架将这些图表示与轻量级(<4B)本地LLM结合,统一拓扑特征与语义推理,同时避免大型云模型的成本和隐私问题。在Java漏洞检测(二分类)上评估,我们的方法达到93.57%准确率——比基于图注意力网络的嵌入高8.36%,比Qwen2.5 Coder 3B等预训练LLM基线高17.81%。除准确率外,该方法提取显著子图并生成自然语言解释,提升对开发者的可解释性。这些结果为可扩展、可解释且本地部署的工具铺平道路,可将漏洞分析从纯语法检查转向更深层的结构和语义洞察,促进在现实安全软件开发中的更广泛采用。
英文摘要:
Software vulnerabilities remain a persistent risk, yet static and dynamic analyses often overlook structural dependencies that shape insecure behaviors. Viewing programs as heterogeneous graphs, we capture control- and data-flow relations as complex interaction networks. Our hybrid framework combines these graph representations with light-weight (<4B) local LLMs, uniting topological features with semantic reasoning while avoiding the cost and privacy concerns of large cloud models. Evaluated on Java vulnerability detection (binary classification), our method achieves 93.57% accuracy-an 8.36% gain over Graph Attention Network-based embeddings and 17.81% over pretrained LLM baselines such as Qwen2.5 Coder 3B. Beyond accuracy, the approach extracts salient subgraphs and generates natural language explanations, improving interpretability for developers. These results pave the way for scalable, explainable, and locally deployable tools that can shift vulnerability analysis from purely syntactic checks to deeper structural and semantic insights, facilitating broader adoption in real-world secure software development.