发表机构
University of Electronic Science and Technology of China; Chinese Academy of Sciences(电子科技大学; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RAG在动态对抗环境中面临的冲突证据挑战,提出无训练框架EvoTrustRAG,通过构建冲突证据图并利用时间关系等进行归因,提升了冲突归因的准确率与宏F1值,降低了协同攻击下的错误率。
AI 中文摘要
检索增强生成(Retrieval-Augmented Generation,RAG)借助外部知识提升大语言模型的事实性,但在动态与对抗环境中,冲突证据仍是核心挑战。现有方法常将冲突视为静态不一致性,选择更可靠的知识,却忽略同一冲突可能源于合法的知识演化、恶意操纵或未解决的不确定性。我们将冲突起源归因定义为RAG中的新问题:识别冲突证据的哪种解释由可观测上下文支持,而非仅判断应信任哪个事实。我们提出EvoTrustRAG,这是一种在答案生成前用于感知演化的冲突归因与证据处理的无训练框架。EvoTrustRAG将基于片段的检索事实表示为冲突证据图,利用时间关系、支持结构和辅助一致性评估基于片段的演化与定向干预假设,并将局部决策投射到每个冲突组的全局一致解释中。归因过程会确定是保留前后状态作为时间知识、将干预候选与主要上下文分离,还是保留未解决的冲突供生成器处理。与专注于事后分析的溯源方法不同,EvoTrustRAG在推理过程中判断冲突证据是否遵循合理的知识演化、表现出类似干预的支持,或无法可靠归因。实验表明,EvoTrustRAG在基准原生冲突设置上达到81.4%的平均准确率,将归因宏F1值从最强基线的72.2%提升至79.1%,并在最强协同攻击下将错误率从31.2%降至16.0%。
英文摘要
Retrieval-Augmented Generation (RAG) improves the factuality of large language models with external knowledge, yet conflicting evidence remains a fundamental challenge in dynamic and adversarial environments. Existing approaches often treat conflicts as static inconsistencies and select more reliable knowledge, overlooking that the same conflict may arise from legitimate knowledge evolution, malicious manipulation, or unresolved uncertainty. We formulate conflict origin attribution as a new problem in RAG: identifying which explanation of conflicting evidence is supported by observable context rather than simply which fact should be trusted. We propose EvoTrustRAG, a training-free framework for evolution-aware conflict attribution and evidence handling before answer generation. EvoTrustRAG represents span-grounded retrieved facts as a conflict evidence graph, evaluates grounded evolution and directional intervention hypotheses using temporal relations, support structure, and auxiliary consistency, and projects local decisions onto a globally consistent explanation of each conflict group. The attribution determines whether earlier and later states are preserved as temporal knowledge, an intervention candidate is separated from the primary context, or an unresolved conflict remains visible to the generator. Unlike provenance-based approaches focused on post-hoc analysis, EvoTrustRAG determines during inference whether conflicting evidence follows plausible knowledge evolution, exhibits intervention-like support, or cannot be reliably attributed. Experiments show that EvoTrustRAG achieves 81.4% average accuracy on benchmark-native conflict settings, improves attribution macro-F1 from 72.2% to 79.1% over the strongest baseline, and reduces the error rate under the strongest coordinated attack from 31.2% to 16.0%.