arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16777cs.CL

通过多轮对话说服基准测试LLMs的事实鲁棒性

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

Zhuoang Cai

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLMs事实鲁棒性,提出SAST-IR框架模拟无状态目标对抗说服攻击,发现简单策略成功率高达96%,并揭示复杂性悖论,即简单策略更易实现真正说服。

中文摘要 AI 辅助

随着大型语言模型(LLMs)日益成为主要的知识检索界面,它们对抗“说服攻击”(即试图注入错误信息或强制执行反事实的尝试)的鲁棒性已成为一个关键的安全问题。现有的红队测试框架通常在多轮对话中评估模型,其中目标模型保留完整的对话历史。我们识别出这种设置中的一个关键缺陷,称为“拒绝惯性”:模型最初的拒绝往往会在后续轮次中传播,主要是为了保持上下文一致性,从而掩盖了其对复杂、孤立说服尝试的真实脆弱性。为了严格评估最先进(SOTA)模型的“冷启动”防御能力,我们引入了SAST-IR(有状态攻击者,无状态目标-迭代细化)框架。通过对目标强制执行记忆清除,同时保留攻击者的历史,我们使用多轮(无状态)迭代模拟了最坏情况的对抗性设置。利用CP-Agent(认知说服代理),一个增强的诊断引导代理,我们在自定义的CounterFact-Strict数据集(N=50)上的实验得出了令人震惊的结果:简单、多样的攻击策略达到了惊人的96%成功率,暴露了无记忆防御的严重脆弱性。此外,我们揭示了一个“复杂性悖论”:虽然复杂、迭代细化的攻击是有效的,但它们往往触发防御性顺从,而简单策略实现了更高的真正说服率(84.7%)。我们的代码和数据集可在GitHub上获取,此https URL。

英文摘要

As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a critical safety concern. Existing red-teaming frameworks typically evaluate models in multi-turn dialogues where the target model retains full conversation history. We identify a critical flaw in this setting termed \textbf{``Refusal Inertia''}: a model's initial refusal often propagates through subsequent turns largely to maintain contextual consistency, thereby masking its true vulnerability to sophisticated, isolated persuasion attempts. To rigorously evaluate the ``cold-start'' defense capabilities of SOTA models, we introduce the \textbf{SAST-IR} (Stateful Attacker, Stateless Target - Iterative Refinement) framework. By enforcing a memory wipe on the target while retaining the attacker's history, we simulate a worst-case adversarial setting using \textbf{multi-turn} (stateless) iterations. Leveraging \textbf{CP-Agent} (Cognitive Persuasion Agent), an enhanced diagnosis-guided agent, our experiments on the custom \textsc{CounterFact-Strict} dataset ($N=50$) yield alarming results: simple, diverse attack strategies achieved a staggering \textbf{96\%} success rate, exposing severe brittleness in memory-less defense. Furthermore, we reveal a \textbf{``Complexity Paradox''}: while complex, iteratively refined attacks are effective, they often trigger defensive compliance, whereas simple strategies achieve a higher rate of genuine persuasion (\textbf{84.7\%}). Our code and dataset are available at GitHub, https://github.com/cza1006/llm-persuasion-defense.

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑