时间拉锯战:通过轨迹方差可视化和检测扩散模型中的RAG冲突
The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance
AI总结:
针对RAG在扩散模型中的知识冲突,提出轨迹方差分数(TVS)作为可解释检测指标,仅需两次并行推理,在多个数据集上以逻辑回归实现高准确率,并验证跨架构可迁移性。
AI中文摘要:
检索增强生成(RAG)在离散扩散语言模型中引入了一种特定的失败模式:当检索到的上下文与参数化知识相矛盾时,迭代去噪过程成为竞争知识来源之间可见的战场。我们识别出时间语义发散作为检测这些冲突的可观测指标,并引入了轨迹方差分数(TVS),这是一种对这种发散进行简单且可解释的度量。TVS计算独立随机去噪轨迹中答案嵌入的平均成对余弦距离,捕捉参数化吸引子与上下文吸引子之间的时间拉锯战。TVS仅需最少两次并行推理运行,计算开销轻量。在四个不同的数据集(Synthetic、SciQ、PopQA和CounterFact)上,使用TVS的简单逻辑回归分类器在LLaDA上达到了70.10%的准确率和0.7647的AUROC。在Dream 7B上,将轨迹数量从两条增加到五条,准确率从63.91%提升到69.62%。更复杂的序列模型相对于线性分类器仅提供边际改进。在LLaDA和Dream 7B上的评估表明,冲突引起的轨迹动态及其关键属性可跨不同的扩散架构转移。
英文摘要:
Retrieval-Augmented Generation (RAG) introduces a specific failure mode in discrete diffusion language models: when retrieved context contradicts parametric knowledge, the iterative denoising process becomes a visible battleground between competing knowledge sources. We identify temporal semantic divergence as an observable for detecting these conflicts and introduce the Trajectory Variance Score (TVS), a simple and interpretable measure of this divergence. TVS computes the mean pairwise cosine distance of answer embeddings across independent stochastic denoising trajectories, capturing the temporal tug of war between parametric and contextual attractors. Requiring as few as two parallel inference runs, TVS is computationally lightweight. Across four diverse datasets (Synthetic, SciQ, PopQA, and CounterFact), a simple Logistic Regression classifier using TVS achieves $70.10\%$ accuracy and $0.7647$ AUROC on LLaDA. On Dream 7B, increasing the number of trajectories from two to five improves accuracy from $63.91\%$ to $69.62\%$. More complex sequential models provide only marginal improvements over the linear classifier. Evaluation across LLaDA and Dream 7B demonstrates that conflict-induced trajectory dynamics and their key properties transfer across distinct diffusion architectures.