arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10627cs.CL

分解诱导的上下文-记忆冲突:事实核查管道与其自身源文本的矛盾

Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

Yu-Feng Yen

AI总结:

该研究发现分解后验证的事实核查管道存在分解诱导的上下文-记忆冲突(DI-CC),线性探针可检测该问题,经典缓解方法可迁移但代价大,确立了其为需关注的失败模式。

AI中文摘要:

分解后验证的管道,包括FActScore式事实核查器和长文本事实性评估器,会先将段落拆分为原子主张,再对每个主张进行核查,分解过程本身被视为中性预处理步骤。我们证明这一假设不成立:分解器可被诱导用自身参数化信念替代源段落内容,生成与应忠实总结的文本相矛盾的主张,我们将此称为分解诱导的上下文-记忆冲突(DI-CC),并表明其机制与经典上下文-记忆冲突相同,只是发生在与 prior work 所研究不同的管道阶段。仅在经典上下文-记忆冲突数据(NQ-Swap)上训练的线性探针,从未接触过任何分解输出,可显著区分产生DI-CC的分解位置与忠实分解(AUC=0.86-0.88,置换检验p<0.0005)。现有无参考基线SelfCheckGPT式自一致性采样完全无法检测DI-CC(AUC=0.51,随机水平),因为DI-CC内容稳定可复现,在重采样中反复出现,与自一致性方法依赖的变异性不同。来自经典设置的无训练缓解方法上下文感知解码可迁移至分解并抑制DI-CC,但代价严重:在指代密集条件下,许多分解无法解析,常因分解器编造不同实体。我们认为此缓解方法尚未可部署。我们进一步表征该机制的边界:其自然出现率过低,未在自然出现的幻觉文本中显现,且需要最小模型规模才能检测。我们确立DI-CC为一种真实、有机制基础且部分可处理的失败模式,其范围不应被夸大。

英文摘要:

Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims before checking each one. Decomposition itself is treated as a neutral preprocessing step. We show it is not: a decomposer can be induced to substitute its own parametric belief for what the source passage says, producing a claim that contradicts the text it was supposed to summarize faithfully. We call this Decomposition-Induced Context-Memory Conflict (DI-CC) and show it is mechanistically the same phenomenon as classical context-memory conflict, occurring inside a different pipeline stage than prior work has examined. A linear probe trained only on classical context-memory conflict data (NQ-Swap), never exposed to any decomposition output, significantly separates decomposition positions that produce DI-CC from faithful decompositions (AUC = 0.86-0.88, permutation p < 0.0005). An existing reference-free baseline, SelfCheckGPT-style self-consistency sampling, fails to detect DI-CC at all (AUC 0.51, chance-level), because DI-CC content is stably recoverable and recurs across resamples, unlike the variability self-consistency methods rely on. Context-aware decoding, a training-free mitigation from the classical setting, transfers to decomposition and suppresses DI-CC, but at a severe cost: many decompositions under coreference-heavy conditions fail to parse, often because the decomposer fabricates a different identity. We do not consider this mitigation deployment-ready. We further characterize the mechanism's boundaries: its natural occurrence rate is too sparss not manifest on naturally-occurring hallucinatedtext, and it requires a minimum model scale to detecablish DI-CC as a real, mechanistically grounded, andpartially treatable failure mode, with a scope we chhan overstate.

补充信息

↑