AI 中文总结
研究针对静态漏洞分析的局限,提出TaintRadar方法,通过增强代码属性图的三个语义分析层改进污点分析,在合成基准和实际系统评估中减少误报、提高准确率,重新发现多数已知CVE并找出多个零日漏洞,提升了静态污点分析的多方面性能。
AI 中文摘要
尽管取得了长足进步,但静态漏洞分析仍存在三个关键限制:粗粒度清理建模,将验证视为二元屏障;数据库盲区,破坏跨持久层的污点跟踪;浅层面向对象分析,错过字段级和过程间数据流。这些缺陷源于一个共同根源:基于代码属性图(CPG)的污点分析缺乏联合建模清理、持久化和对象别名的组合语义层。因此,现有工具会产生过多误报或完全错过关键攻击路径。为解决这些限制,我们提出了TaintRadar,一种通过三个语义分析层系统增强CPG的方法。首先,漏洞类型清理使用传递函数和上下文敏感参数绑定计算节点级安全保证。其次,持久化感知传播集成数据库模式约束和查询安全分析,以跟踪穿越共享数据库状态的多脚本攻击路径。最后,对象感知可达定义结合调用和有界变量别名上下文,精确建模跨方法边界的对象字段突变。我们在合成基准和实际系统上评估了TaintRadar。在SARD基准上,TaintRadar大幅减少误报,同时保持80%的总体准确率。在19个实际PHP应用程序中部署时,它重新发现了大多数已知的CVE,并发现了29个已确认的零日漏洞,包括26个SQL注入和3个存储XSS漏洞,这些漏洞已获得CVE标识符。这些结果表明,语义感知图增强显著提高了静态污点分析的精度、覆盖率和实际效用。
英文摘要
Despite significant advances, static vulnerability analysis suffers from three critical limitations: coarse sanitization modeling, which treats validation as a binary barrier; database blindness, which breaks taint tracking across persistence layers; and shallow object-oriented analysis, which misses field-level and interprocedural data flows. These flaws stem from a common root cause: Code Property Graph (CPG)-based taint analyses lack a compositional semantic layer to jointly model sanitization, persistence, and object aliasing. Consequently, existing tools generate excessive false positives or miss critical attack paths entirely. To address these limitations, we present TaintRadar, an approach that systematically augments CPGs with three semantic analysis layers. First, vulnerability-typed sanitization computes node-level safety guarantees using transfer functions and context-sensitive parameter binding. Second, persistence-aware propagation integrates database schema constraints and query safety analysis to track multi-script attack paths traversing shared database states. Finally, object-aware reaching definitions combine calling and bounded variable alias contexts to precisely model object-field mutations across method boundaries. We evaluate TaintRadar on both synthetic benchmarks and real-world systems. On the SARD benchmark, TaintRadar drastically reduces false positives while maintaining 80% overall accuracy. Deployed across 19 real-world PHP applications, it rediscovered the majority of known CVEs and uncovered 29 confirmed zero-day vulnerabilities, including 26 SQL injection and 3 stored XSS vulnerabilities, that have already received CVE identifiers. These results demonstrate that semantic-aware graph augmentation significantly improves the precision, coverage, and practical utility of static taint analysis.
Comments12 pages, 3 figures, 5 tables, Under Review