漏洞评估中的AI垃圾内容与幻觉:关于推理失败及可信缓解措施的综述
AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation
浏览论文内容
中文总结 AI 辅助
本综述针对LLMs用于漏洞评估时产生的AI垃圾内容与幻觉问题,分析其源于专家演绎推理与LLMs自回归生成的差距,提出神经符号验证等缓解方案及CVE-Bench等评估工具,为AI驱动的安全分诊系统提供路线图。
中文摘要 AI 辅助
大型语言模型(LLMs)在网络安全领域的应用变革了漏洞评估,但也因“AI垃圾内容”的无节制扩散引发了可信性危机。这类人工制品包括幻觉漏洞、看似合理却不正确的补丁、语义重新包装的漏洞报告,给人工分诊流程带来了类似拒绝服务攻击的认知负担。本文综述实证证据,识别统一机制并探寻可信分诊的路径。我们通过结构化文献综述构建AI垃圾内容的分类体系,剖析其根本原因:安全专家的因果演绎推理与当前LLMs的自回归概率生成之间存在差距。我们通过可测量的代理指标——演绎覆盖率(Deductive Coverage Score)来量化这一差距,结果显示思维链提示和使用工具的智能体可缩小该差距但无法完全消除。我们综述缓解策略,指出被动检测和水印技术针对的是来源而非正确性,面临基本的熵约束。我们转而倡导主动神经符号验证,将每个分诊流程组件映射到对安全输入有明确限制的现有系统。最后,我们明确了两个评估工具:CVE-Bench和Slop-Score,包括数据集构建、指标公式和反博弈条款。通过将评估从语言流畅性转向数学可验证性,本综述为保障新兴AI驱动的分诊系统提供了路线图。
英文摘要
The integration of Large Language Models (LLMs) into cybersecurity has transformed vulnerability assessment, but it has also produced a trustworthiness crisis driven by the unchecked proliferation of "AI slop." These artifacts, hallucinated vulnerabilities, plausible but incorrect patches, and semantically repackaged bug reports, impose a cognitive burden on human triage pipelines that mirrors a denial-of-service attack. This paper surveys the empirical evidence, identifies a unifying mechanism, and traces a path toward trustworthy triage. We formalize a taxonomy of AI slop grounded in a structured literature review and dissect its root cause: the gap between the causal deductive reasoning of security experts and the autoregressive probabilistic generation of current LLMs. We operationalize this gap through a measurable proxy, the Deductive Coverage Score, and show that chain-of-thought prompting and tool-using agents narrow but do not close it. We review mitigation strategies and argue that passive detection and watermarking target provenance rather than correctness, facing fundamental entropy constraints. We instead advocate for active neuro-symbolic verification, mapping each pipeline component to prior systems with documented limits on security inputs. Finally, we specify two evaluation instruments, CVE-Bench and Slop-Score, including dataset construction, metric formulas, and anti-gaming provisions. By shifting evaluation from linguistic fluency to mathematical verifiability, this survey provides a roadmap for securing emerging AI-driven triage systems.
发表机构
- University of New South Wales(新南威尔士大学)
- University of Wollongong(伍伦贡大学)
- Nanjing University of Information Science & Technology(南京信息工程大学)
- Griffith University(格里菲斯大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。