发表机构
Imperial College London; Safe Intelligence(帝国理工学院; 安全智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出新框架,通过语义等价对抗攻击测试LLMs鲁棒性,发现先进模型易受语义保留扰动影响致忠实度大幅下降,凸显需开发鲁棒接地的LLM架构与训练目标。
AI 中文摘要
大语言模型(LLMs)常与外部知识源结合以提升事实准确性并减少幻觉,例如检索增强生成(RAG)方法,但这些系统仍易受内在幻觉影响,即模型生成未被检索证据支持的不忠实或虚构信息。本文提出一种新框架,通过使用经对抗优化方法得到的用户查询的自然语义等价变体进行压力测试,以评估模型对该现象的鲁棒性。我们将该框架(强制严格语义等价约束和内在幻觉目标)应用于白盒、灰盒、黑盒对抗设置下的多种对抗攻击技术,在3个数据集上对5个开源生成模型和5个闭源生成模型评估这些攻击,结果显示即便最先进的模型也极易受语义保留扰动影响,此类扰动会显著降低上下文忠实度(GPT-5-mini的忠实度降幅达50%)。研究发现,即便是最先进的LLMs,对上下文证据的忠实使用仍很脆弱,这促使人们开发能强制实现与表面查询形式无关的鲁棒接地的架构和训练目标。代码可在this https URL获取
英文摘要
Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG). However, these systems remain susceptible to intrinsic hallucinations, where the model generates unfaithful or fabricated information that is not supported by the retrieved evidence. We propose a novel framework to assess model robustness against this phenomenon by stress-testing using natural, semantically equivalent variations of a user query found via adversarial optimization methods. We apply our framework, which enforces strict semantic equivalence constraints and an intrinsic hallucination objective, to a range of adversarial attack techniques across white-box, gray-box, and black-box adversarial settings. Evaluating these attacks on 5 open-source and 5 closed-source generator models across 3 datasets, we demonstrate that even state-of-the-art models are highly susceptible to meaning-preserving perturbations, which significantly degrade contextual faithfulness (by up to 50% for GPT-5-mini). Our findings indicate that faithful use of in-context evidence remains fragile even in state-of-the-art LLMs, motivating architectures and training objectives that enforce robust grounding independent of surface query form. Code is available at: https://github.com/atriviveksharma/intrinsic_hall
CommentsTo be presented at COLM 2026