arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扫描器-LLM级联中的案例级验证:克服告警聚合瓶颈以扩展FRR-TPR权衡空间

Case-Level Verification in Scanner-LLM Cascades: Overcoming the Alert Aggregation Bottleneck to Expand the FRR-TPR Trade-off Space

Hao Sun, Yibin Yao, Chaohai Xie, Yuqun Lin

arXiv 2610.08406首次发表:更新:

发表机构

Shenzhen Information Security Management Center; Shenzhen Secidea Network Security Technology Co. Ltd(深圳市信息安全管理中心; 深圳市思科达网络安全科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对DAST扫描器与LLM级联中告警聚合导致的FRR-TPR权衡瓶颈,提出案例级逐案例验证策略,将决策粒度从告警级降至实例级,在DVWA和WebGoat上实现FRR提升4.7个百分点,TPR提升6.9个百分点。

AI 中文摘要

动态应用安全测试(DAST)扫描器具有高召回率,但也会产生大量误报,导致大量人工分诊成本。大型语言模型(LLMs)在用于独立检测时,虽然实现了极高的召回率(95.4%-100%),但其误报率也高得令人望而却步(49.6%-85.0%),因此无法作为扫描器的独立替代品。一种自然的解决方案是采用两阶段级联,即先由扫描器检测,再由LLM进行验证。然而,一个在实践中长期被忽视的验证粒度问题造成了结构性瓶颈:告警聚合将多个真案例和假案例绑定到一个共享决策单元中,使得移除一个误报不可避免地会消除同一告警组内聚合的真阳性。这在误报降低率(FRR)和真阳性率(TPR)之间造成了权衡瓶颈。我们通过证明告警级误报集合是案例级误报集合的子集来形式化这一瓶颈,并提出了一种案例级、逐案例验证策略,将决策粒度从告警级别转移到实例级别,独立重放HTTP请求并对每个检测到的案例做出独立判定。在Damn Vulnerable Web Application(DVWA)和WebGoat双测试平台上的评估表明,经验上最佳的告警级工作点实现了FRR=42.86%(TPR=51.7%)。案例级基线实现了FRR=47.6%,提升了4.7个百分点(+4.7 pp),而案例聚焦证据验证提示(CEV-Prompt)在相同FRR下将TPR从55.2%提高到62.1%。

英文摘要

Dynamic Application Security Testing (DAST) scanners achieve high recall but also produce a large number of false positives, resulting in substantial manual triage costs. Large Language Models (LLMs), when used for independent detection, achieve extremely high recall (95.4%-100%) but also exhibit prohibitively high false positive rates (49.6%-85.0%), precluding their use as standalone replacements for scanners. A natural solution is a two-stage cascade consisting of scanner detection followed by LLM verification. However, a verification-granularity issue that has long been overlooked in practice creates a structural bottleneck: alert aggregation binds multiple true and false cases into a shared decision unit, such that removing a false positive inevitably eliminates true positives aggregated within the same alert group. This creates a trade-off bottleneck between the False-positive Reduction Rate (FRR) and the True-positive Rate (TPR). We formalize this bottleneck by showing that the alert-level false-positive set is a subset of the case-level false-positive set, and introduce a Case-Level, per-case verification strategy that shifts the decision granularity from the alert level to the instance level, independently replaying HTTP requests and making an independent determination for each detected case. Evaluation on the dual testbeds of Damn Vulnerable Web Application (DVWA) and WebGoat shows that the empirically best Alert-Level operating point achieves FRR=42.86% (TPR=51.7%). The Case-Level Baseline achieves FRR=47.6%, an improvement of 4.7 percentage points (+4.7 pp), while the Case-Focused Evidence Verification Prompt (CEV-Prompt) increases TPR from 55.2% to 62.1% at the same FRR.

CommentsAccepted at the 3rd International Conference on Intelligent Computing and Data Analysis (ICDA 2026). To appear in the ACM International Conference Proceedings Series (ICPS). 17 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑