arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型增强的语义感知类型检查的漏洞发现

Finding Vulnerabilities via LLM-Augmented Semantics-Aware Type-Checking

Ruizhe Wang, Meng Xu, N. Asokan

arXiv 2608.14533首次发表:更新:

AI 中文总结

本文提出语义感知类型系统SETYPE,由LLMs执行类型推断与检查,构建PYSETYPE原型,在真实Python Web应用中实现87%精度、88%准确率,识别出15个潜在零日漏洞,9个获开发者确认。

AI 中文摘要

传统静态分析的漏洞检测依赖安全专家将不安全编码模式编码为算法规则,但该方法常关注句法模式,忽略代码中更深层的语义信息,如变量和函数名的含义。随着软件系统日益复杂,仅用句法规则建模漏洞的难度不断增加。本文提出一种语义感知的软件漏洞检测方法,给出SETYPE,一种仅基于源代码中符号和表达式的自然语言含义直接从源代码推导的语义感知类型系统。在SETYPE类型系统中,类型推断和检查均由大语言模型(LLMs)执行,类型检查失败则表明存在潜在漏洞。我们构建了PYSETYPE原型以验证SETYPE检测Python Web应用漏洞的可行性,对真实应用的评估显示其检测精度达87%,检测准确率达88%。通过PYSETYPE,我们识别出15个潜在零日漏洞,其中9个已被开发者确认。

英文摘要

Vulnerability detection via static analysis traditionally relies on security experts encoding insecure coding patterns into algorithmic rules. However, this approach often focuses on syntactic patterns and overlooks deeper semantic information in the code, such as the meanings of variable and function names. As software systems grow more complex, modeling vulnerabilities using only syntactic rules becomes increasingly challenging. In this paper, we propose a semantics-aware approach to detecting software vulnerabilities. We present SETYPE, a semantics-aware type system that can be derived directly from source code based solely on the meanings of symbols and expressions in natural language. In the SETYPE type system, both type inference and checking are performed by Large Language Models (LLMs), and a failed type check indicates a potential vulnerability. We prototype PYSETYPE to demonstrate the feasibility of SETYPE for detecting vulnerabilities in Python web applications. Our evaluation on real-world applications achieves 87% detection precision and 88% detection accuracy. Using PYSETYPE, we identified 15 potential zero-day vulnerabilities, nine of which were confirmed by developers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑