AI 中文总结
本研究提出CAUSEC因果分析框架,通过对57个SAST假设的定性分析及在4款工具上的测试,揭示SAST性能与假设间的因果关系,为未来研究提供3个要点。
AI 中文摘要
静态应用安全测试(SAST)工具在工业界和学术界被广泛使用,这类工具通常会做出牺牲检测能力以换取更高性能的设计选择,即提升精度、缩短运行时间或增强可扩展性。这些设计选择依赖于针对目标代码或分析技术本身的某些假设,因此这些假设会通过其影响的设计选择直接作用于检测结果。这引出了一个关键问题:检测能力的牺牲是否真的帮助工具实现了预期的性能提升?也就是,这些底层假设是否有效?本文基于一个关键观察——这些工具做出的假设通常具有因果性质,来尝试解决该问题。我们提出CAUSEC,一个因果分析框架,该框架使SAST假设可被测试,并能解释在特定假设下性能变化的原因,而非仅简单的相关性。CAUSEC将SAST工具的假设形式化为安全假设的抽象概念,结合假设驱动的因果建模、效应估计与验证,以测试其有效性并研究影响它的因素。为理解安全假设通常包含的内容,我们对检测加密API误用的SAST工具进行了系统文献综述,进而发现并定性分析了57个假设。随后,我们使用包含57038个警报的手动标注基准数据集,通过在四个高度相关的工具中测试一个流行假设,证明了CAUSEC的实用性和稳健性。我们的分析得出了若干关于假设和因果效应的关键发现,并将其提炼为对未来研究的3个要点。
英文摘要
Static Application Security Testing (SAST) tools are widely used in both industry and academia. Such tools often make design choices that sacrifice detection to achieve higher performance, i.e., increased precision, decreased runtime, or increased scalability. These design choices rely on certain assumptions regarding the target code or the analysis technique itself. Hence, the assumptions directly impact the detection outcome through the design choices they influence. This motivates a key question: do the sacrifices in the detection capabilities actually help tools achieve the expected performance gains? That is, are the underlying assumptions valid? This paper seeks to address this question by relying on a key observation that the assumptions made by these tools are generally of a causal nature. We propose CAUSEC, a causal analysis framework that makes SAST assumptions testable and explains why the performance changes given certain assumptions, beyond simple correlations. CAUSEC formalizes the assumptions of the SAST tool into the abstraction of a security assumption and combines assumption-driven causal modeling with effect estimation and validation to test its validity and investigate the factors affecting it. To understand what security assumptions generally entail, we perform a systematic literature review of SASTs that detect crypto-API misuse, leading to the discovery and qualitative analysis of 57 assumptions. We then demonstrate the utility and robustness of CAUSEC by testing a popular assumption in four highly relevant tools, using a manually labeled ground truth dataset consisting of 57,038 alerts. Our analysis leads to several key findings that represent insights regarding assumptions and causal effects, which we distill into 3 takeaways for future work.