发表机构
Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RAG系统面临的知识篡改攻击问题,提出拜占庭容错的安全协同RAG框架SecureCollaRAG,通过多源知识验证机制与动态GNN可信度评分实现攻击防护,在非IID数据下保持鲁棒性。
AI 中文摘要
尽管检索增强生成系统在一定程度上解决了大语言模型的幻觉问题,但也引入了知识篡改攻击的新漏洞。攻击者利用这些漏洞,通过投毒RAG系统提供的文档来操纵大语言模型的输出。为应对这一威胁,我们提出了SecureCollaRAG,一种利用多源知识验证机制的拜占庭容错协同RAG框架。我们的方法使智能体系统能够通过基于动态图神经网络(GNN)的可信度评分安全验证文档来源,有效防止隐蔽的知识篡改攻击,同时保持必要的领域知识完整性。通过广泛的评估和形式化分析,我们证明SecureCollaRAG在非独立同分布(non-IID)数据分布下对攻击者保持鲁棒性。
英文摘要
While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions.
Journal refProceedings of the ACM Web Conference 2026, pages 2661-2672, 2026