arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可信检索增强生成:用于检测生成式人工智能系统中错误信息和知识投毒的评估智能体

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson

arXiv 2608.21095首次发表:更新:

发表机构

Faculty of Information Technology and Communication Sciences, Tampere University(坦佩雷大学信息技术与传播科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对RAG系统的安全-可靠性鸿沟,提出结合NLI验证、五信号投毒检测器及可信指数的评估智能体,在TruthfulQA等数据集上实现高检测性能,可阻止RAG系统的不安全指令注入。

AI 中文摘要

检索增强生成(Retrieval-Augmented Generation, RAG)将大型语言模型(Large Language Model, LLM)的输出建立在外部知识基础上,但RAG系统通常信任所检索到的任何内容,从而产生安全-可靠性鸿沟:高语义相关性并不保证事实真实性。攻击者利用这一点实施知识投毒,插入恶意文档以引发定向错误信息。我们提出一种评估智能体,这一中间件结合了自然语言推理(Natural Language Inference, NLI)事实验证、具有相关性加权聚合的五信号投毒检测器,以及针对高污染场景的带非线性抑制项的可信指数T = 0.4F + 0.35C + 0.25(1 - P)。在使用Llama 3.3 70B的TruthfulQA数据集上,该智能体达到91%的准确率和100%的精确率,对指令注入的召回率为100%,而原位编辑(如实体替换)仍难以检测。在三种LLM上,可信指数保持区分性,受试者工作特征曲线下面积(Receiver Operating Characteristic Area Under the Curve, ROC-AUC)为0.73至0.81;生成风格比模型规模更重要,针对每个LLM的阈值校准可恢复基线竞争准确率,而较弱的FEVER结果表明,跨数据集泛化需要特定领域校准。在软件工程用例中,基于开放全球应用安全项目(Open Worldwide Application Security Project, OWASP)十大和常见弱点枚举(Common Weakness Enumeration, CWE)指导的安全编码助手中,该智能体可靠地阻止不安全建议的指令注入(F1值为92%),而矛盾和微妙语义弱化仍难以检测。总体而言,该智能体在生成前测量对投毒上下文的检测,而非LLM是否采用注入的错误信息。我们在链接this https URL发布了所提出的方法、攻击生成器和实验产物。

英文摘要

Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not guarantee factual truth. Adversaries exploit this through knowledge poisoning, inserting malicious documents to cause targeted misinformation. We propose an Evaluation Agent, middleware that combines Natural Language Inference (NLI) factual verification, a five-signal poison detector with relevance-weighted aggregation, and a Trust Index T = 0.4 F + 0.35 C + 0.25 (1 - P ) with a non-linear dampener for high-contamination contexts. On TruthfulQA with Llama 3.3 70B, the agent reaches 91% accuracy and 100% precision, with 100% recall on instruction injection, while in-place edits, such as entity swaps, remain hard to detect. Across three LLMs the Trust Index stays discriminative, with a Receiver Operating Characteristic Area Under the Curve (ROC-AUC) of 0.73 to 0.81; generation style matters more than model size, and per-LLM threshold calibration restores baseline competitive accuracy, whereas a weaker FEVER result shows that cross-dataset generalization requires domain-specific calibration. In a software-engineering use case, a secure-coding assistant over guidance from the Open Worldwide Application Security Project (OWASP) Top 10 and the Common Weakness Enumeration (CWE), the agent reliably blocks instruction injection of unsafe advice (F1 92%), while contradiction and subtle semantic weakening remain hard. Throughout, the agent measures detection of poisoned context before generation, not whether the LLM adopts the injected misinformation. We release the proposed approach, attack generator, and experimental artifacts at the link: https://github.com/GPT-Laboratory/TrustworthyRAG.

Comments7 pages, 1 figure. Accepted for publication in the Main Research Track of the Twenty-First International Conference on Software Engineering Advances (ICSEA 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑