arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35308cs.CL

多轮LLM污染中的认知策略分歧:一种协议梯度研究

Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation

Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出五种污染协议,揭示多轮LLM中错误前提采纳的机制差异,并发现不同模型在权威遵从与指令服从上存在显著分歧,强调对话历史需来源感知设计。

中文摘要 AI 辅助

大型语言模型将对话历史视为未经验证的上下文:先前轮次中注入的错误前提可能被当作事实采纳,我们将这种失败模式称为会话级污染。我们引入了五种沿来源权威梯度排列的污染协议,在保持错误前提不变的情况下隔离不同的失败机制,并在温度为零的条件下(22,500轮次)跨十个知识领域评估了GPT-5.4 Mini、Gemini-3.1 Flash-Lite和GLM-4.5-Air,使用经人工金标准验证的双轨自动评判器(Cohen's κ = 0.901)。GPT-5.4 Mini在所有500个会话中表现出零采纳,这是一种会话级别的与内容无关的策略;令牌级探测显示其底层边际虽然较大,但有限。Gemini-3.1 Flash-Lite遵循陡峭的权威梯度:对自我归因的错误陈述采纳率为0.1%,对用户引用的来源为23.5%,对系统注入的权威为68.2%,在指令覆盖下达到94.0%。GLM-4.5-Air表现出较浅的梯度(15.8%对84.2%),68个百分点的分离证实了权威遵从和指令服从在同一架构中是不同的机制。恢复情况也存在分歧:GLM在94.5%受影响的会话中恢复,而Gemini受影响的会话中有26.1%从未恢复,在指令覆盖下上升至40.0%。对话历史是一个不受信任的攻击面,需要具有来源感知的系统设计;完整框架已作为开源基准发布。

英文摘要

Large language models treat conversation history as unverified context, so false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, holding the false premise constant while varying its epistemic framing, and evaluate GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains at temperature zero (22,500 turns), judged by a dual-track automated evaluator validated against a human gold standard (Cohen's kappa = 1.000 for binary adoption; 0.92 linear-weighted for collapse severity). GPT-5.4 Mini recorded zero adoptions across all 500 sessions; a base-model logit probe shows its decision margin is perturbed but large and finite. Gemini-3.1 Flash-Lite followed a steep authority gradient: 0.1% adoption for self-attributed falsehoods, 23.5% for user-cited sources, 68.2% for system-injected authority, and 94.0% under instruction override. GLM-4.5-Air showed a shallower gradient (15.8% vs 84.2%), a 68-point dissociation consistent with authority deference and instruction compliance engaging distinct mechanisms within one architecture. Recovery diverged: GLM recovered in 94.5% of affected sessions, whereas 26.1% of affected Gemini sessions never did, rising to 40.0% under instruction override. Conversation history is an untrusted attack surface requiring provenance-aware system design; the complete evaluation framework is released as an open-source artifact.

发表机构

  • Institut Teknologi Sumatera(苏门答腊理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑