arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对抗智能体系统中的知识篡改:一种拜占庭容错的安全协同RAG框架

Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework

Zhaoqi Wang, Daqing He, Zijian Zhang, Ye Liu, Jiamou Liu, Zhirui Zeng, Zhan Qin, Zhen Li, Xin Li, Hongwei Yao, Jincheng An, Yong Liu, Yi Li, Qi Sun, Xiulei Liu, Liehuang Zhu

arXiv 2608.04366首次发表:更新:

发表机构

Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RAG系统面临的知识篡改攻击问题,提出拜占庭容错的安全协同RAG框架SecureCollaRAG,通过多源知识验证机制与动态GNN可信度评分实现攻击防护,在非IID数据下保持鲁棒性。

AI 中文摘要

尽管检索增强生成系统在一定程度上解决了大语言模型的幻觉问题,但也引入了知识篡改攻击的新漏洞。攻击者利用这些漏洞,通过投毒RAG系统提供的文档来操纵大语言模型的输出。为应对这一威胁,我们提出了SecureCollaRAG,一种利用多源知识验证机制的拜占庭容错协同RAG框架。我们的方法使智能体系统能够通过基于动态图神经网络(GNN)的可信度评分安全验证文档来源,有效防止隐蔽的知识篡改攻击,同时保持必要的领域知识完整性。通过广泛的评估和形式化分析,我们证明SecureCollaRAG在非独立同分布(non-IID)数据分布下对攻击者保持鲁棒性。

英文摘要

While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions.

Journal refProceedings of the ACM Web Conference 2026, pages 2661-2672, 2026

DOI:10.1145/3774904.3792200

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑