arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Memoir:针对静态应用安全测试工具的误报记忆学习、验证与演化

Memoir: Learning, Verifying, and Evolving False-Positive Memories for Static Application Security Testing Tools

Shenyuan Guan, Qiaodan Hou, Yanjun Chen, Xincheng Wen, Jia Feng, Keke Lian, Cuiyun Gao

arXiv 2608.09181首次发表:更新:

AI 中文总结

本研究针对SAST工具的误报问题,提出Memoir框架,通过构建与演化语义记忆实现误报识别,在CWE-Bench-Java上表现优异,且记忆库可跨工具泛化。

AI 中文摘要

静态应用安全测试(SAST)工具已成为现代安全软件开发中不可或缺的工具,但这些工具常产生误报(FP)警报,带来大量人工检查成本并降低开发者信任。现有误报减少方法面临两大核心挑战:一是SAST工具与漏洞类别间的巨大差异,导致难以学习历史误报的重复模式;二是这些方法所用知识基本静态,无法随新验证案例的积累而更新。为解决这些挑战,我们提出Memoir——一种基于记忆的误报识别框架,通过将历史误报警报转化为可复用的语义记忆来实现误报识别。该框架包含两个关键模块:其一,历史语义记忆构建模块,通过大语言模型(LLM)引导的标注、模式聚类与记忆合成,将历史误报警报转化为结构化语义记忆,以捕获可复用的行为模式;其二,记忆驱动的识别与演化模块,在做出最终预测前,会检索相关记忆并针对分类一致性与安全不变量进行语义验证,随后将已验证的预测整合回记忆库,使知识库能随新案例的积累而演化。我们在CWE-Bench-Java上对Memoir进行评估,以证明其在实际安全分析中的有效性。具体而言,Memoir的F1分数达99.43%,召回率为98.88%,精确率为100%,持续优于其他基线方法。此外,我们对某顶尖IT公司生产软件系统开展的工业案例研究显示,所学习的记忆库无需重新训练即可在不同SAST工具间有效泛化。

英文摘要

Static Application Security Testing (SAST) tools have become indispensable in modern secure software devel- opment. However, these tools often generate false-positive (FP) alerts, imposing substantial manual inspection costs and reducing the trust from developers. Existing FP reduction methods still face two primary challenges. First, the large differences among SAST tools and vulnerability categories make it difficult for these methods to learn recurring patterns in historical false positives. Moreover, the knowledge used by these methods are largely static and cannot be updated as newly validated cases accumulate. To address these challenges, we propose Memoir, a memory- driven framework for identifying false positives by transform- ing historical FP alerts into reusable semantic memories. It consists of two key modules. First, historical semantic memory construction converts historical FP alerts into structured semantic memories through LLM-guided annotation, pattern clustering, and memory synthesis to capture reusable behavioral patterns. Moreover, memory-driven identification and evolution retrieves relevant memories and performs semantic verification against taxonomy consistency and security invariants before making the final prediction. It then incorporates verified predictions back into the memory repository, allowing the knowledge base to evolve as new cases accumulate. We evaluate Memoir on CWE- Bench-Java to demonstrate its effectiveness in real-world security analysis. Specifically, Memoir achieves an F1-score of 99.43% with a Recall of 98.88% and perfect Precision, consistently outperforming other baselines. Furthermore, an industrial case study on production software systems from a top IT company shows that the learned memory base generalizes effectively across different SAST tools without retraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑