arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17960cs.CR

COMA:针对安全检索增强生成(Security-RAG)的组合式误导攻击类别及因果反事实防御

COMA: A Compositional Misleading Attack Class on Security-RAG, and a Causal Counterfactual Defense

Chinmay Gondhalekar, Urjitkumar Patel

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对Security-RAG提出COMA组合式误导攻击,含行动破坏与结论翻转两种实现,提出因果反事实防御(ccd)可有效定位并防御该攻击,已发布攻击种子与ccd参考实现。

中文摘要 AI 辅助

安全助手检索到的每一份文档都可能是真实的、无指令的、无矛盾的,但仍可能导致助手正确评估出一个关键的可利用漏洞,却推荐了一个未能修复该漏洞的补救措施。我们研究了安全运营中心中面向分析师的助手所依赖的检索增强生成(RAG)的这一故障,识别出一类攻击——\textit{\textbf{compmis}}(COMA),其中每一份对抗性文档在事实层面都正确、无指令、无矛盾且分布良性,但答案会因它们的\textit{组合}而被误导。我们通过\textit{行动破坏}实现\textit{\textbf{compmis}}:将正确诊断的漏洞引导至较差的补救措施;以及\textit{结论翻转}:通过不可判定的可达性链破坏可利用性结论的稳定性。行动破坏在两个合成领域和一个真实CVE(CVE-2021-33813)上,在每次运行中都能对所有测试的5个模型(包括前沿推理模型)生效;结论翻转则随机生效,随模型能力提升而减弱但从未消失。两者遵循同一原则:当需\textit{推断}而非\textit{读取}消歧事实时,攻击便会成功。我们提出\textit{\textbf{ccd}}(因果反事实防御),这是一种审计方法,用于衡量每份检索文档的留一法因果影响,并标记出其影响集中在低可信度文档上的答案。\textit{\textbf{ccd}}能将攻击定位到攻击者控制的文档,在4个良性多文档对照中无假阳性;自适应影响传播的攻击者则会被\textit{聚合}变体捕获。我们发布了攻击种子和\textit{\textbf{ccd}}参考实现。

英文摘要

Every document a security copilot retrieves can be true, instruction-free, and non-contradictory --- and the copilot can still be driven to assess a critical, exploitable vulnerability correctly and then recommend a remediation that leaves it open. We study this failure in retrieval-augmented generation (RAG) backing analyst-facing copilots in Security Operations Centers, and identify a class of attacks, \emph{\compmis{}} (COMA), in which every adversarial document is factually correct, instruction-free, non-contradictory, and distributionally benign --- yet the answer is misled by their \emph{composition}. We realize \compmis{} through \emph{action-corruption}, which steers a correctly-diagnosed vulnerability toward an inferior remediation, and \emph{verdict-flip}, which destabilizes the exploitability verdict via an undecidable reachability chain. Action-corruption bites all five tested models --- including frontier reasoning models --- on every run, on two synthetic domains and a real CVE (CVE-2021-33813); verdict-flip bites stochastically, decreasing with model capability but never vanishing. A single principle governs both: the attack succeeds when the disambiguating fact must be \emph{inferred} rather than \emph{read}. We propose \ccd{} (Causal Counterfactual Defense), an audit that measures the leave-one-out causal influence of each retrieved document and flags answers whose influence concentrates on low-trust documents. \ccd{} localizes the attack to attacker-controlled documents with no false positives on four benign multi-document controls; an adaptive influence-spreading adversary is caught by an \emph{aggregate} variant. We release attack seeds and a \ccd{} reference implementation.

补充信息

↑