arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

幻觉滚雪球:将多智能体大语言模型流水线中的错误传播建模为状态转移

The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines

Prabhjot Singh, Bhushan Pawar

arXiv 2608.14588首次发表:更新:

AI 中文总结

该研究发现多智能体LLM流水线中存在幻觉滚雪球效应,即幻觉会随状态转移而逃逸,且边界门控验证比末端检查更能有效降低幻觉存活率,还提出了最优验证资源分配策略。

AI 中文摘要

顺序多智能体大语言模型(LLM)流水线会串联专门的智能体,且在交接时未设置验证机制,这一结构缺陷会产生可测量且严重的后果。研究表明,在第1阶段注入的幻觉(hallucinations)不会仅仅持续存在,还会发生转变:原始数值事实会变为推导计算结果,再变为叙述性散文,最终变为经编辑批准的结论。在每一次转变过程中,可检测性会近乎不可逆地下降。研究将此形式化为幻觉滚雪球效应,这是一个包含四个状态(原始事实→推导结果→叙述→不可见)的一阶马尔可夫过程,且通过实验测得各状态边界的逃逸概率分别为24.6%、48.3%和89.3%。在FinanceBench数据集上的4智能体金融分析流水线中,对346个自动注入的幻觉进行测试,结果显示gpt-4o的检测率从第1阶段的72.0%降至第4阶段的50.9%,最终输出中有23.7%的幻觉完全未被检测到。即便测试的最强模型Qwen3.5-397B-A17B在第1阶段的检测率为87.0%,也面临结构上限,预计第4阶段的检测率仅约60%至65%。关键发现是,使用相同检索增强生成(RAG)验证工具的边界门控,与流水线末端检查相比,可将幻觉存活率从58.4%降至16.2%(Cohen's h=-0.911,p<0.000001),而仅进行末端检查相比不进行验证仅提升2.3个百分点。何时进行验证比是否进行验证更重要。该模型可预测n智能体线性流水线的幻觉存活率,并能指导最优验证资源分配:应优先在S₁→S₂阶段投入资源,此时仍有75.4%的幻觉可被捕获,而非在S₃→S₄阶段,此时已有89.3%的幻觉逃逸。

英文摘要

Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persist; they transform: raw numerical facts become derived computations, then narrative prose, then editorially approved conclusions. At each transformation, detectability degrades near-irreversibly. We formalize this as the hallucination snowball effect, a first-order Markov process over four states (Raw Fact $\to$ Derived $\to$ Narrative $\to$ Invisible) with empirically measured per-boundary escape probabilities of 24.6%, 48.3%, and 89.3%. Across 346 automatically injected hallucinations in a 4-agent financial analysis pipeline on FinanceBench, gpt-4o detection drops from 72.0% at Stage 1 to 50.9% at Stage 4, and 23.7% of hallucinations survive completely undetected in the final output. Even the strongest model tested (Qwen3.5-397B-A17B, 87.0% at Stage 1) faces a structural ceiling; projected Stage 4 detection is only ${\sim}$60--65%. Critically, boundary gates using identical RAG verification tools reduce hallucination survival from 58.4% to 16.2% versus end-of-pipeline checking (Cohen's $h = -0.911$, $p < 0.000001$), while end-checking alone achieves merely 2.3 pp improvement over no verification. When you verify matters more than whether you verify. Our model predicts survival for $n$-agent linear pipelines and prescribes optimal verification resource allocation: invest at $S_1{\to}S_2$ first, where 75.4% of hallucinations are still catchable, not at $S_3{\to}S_4$ where 89.3% have already escaped.

Comments10 pages, 3 figures; accepted at the FAGEN Workshop (Failure Modes in Agentic AI), ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑