arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

感知缓慢,抑制迟缓:理解模态在上下文记忆冲突中的作用

Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts

Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick

arXiv 2609.00293首次发表:更新:

发表机构

Brown University(布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究视觉语言模型处理上下文记忆冲突的不对称偏差,发现模型对文本实体偏好上下文信息、对图像实体偏好参数信息,思维链推理无法缩小该差距,增加视觉信息可产生效果,揭示多模态检索增强模型确保一致行为的复杂性。

AI 中文摘要

我们研究视觉语言模型(VLMs)如何处理上下文记忆冲突,即模型接收到的上下文信息与训练期间存储在参数中的信息不一致的情况。我们记录了不对称偏差:模型倾向于偏好文本中出现的实体的上下文信息,但偏好图像中出现的实体的参数信息。我们将这种不对称性与跨模态的后期表示对齐联系起来,表明处理视觉实体所需的更长处理时间会阻碍对模型常规事实召回机制的抑制,从而导致更多参数化答案。思维链推理似乎无法缩小这一差距,但增加上下文中的视觉信息数量确实会产生效果。这些结果表明,随着模型变得越来越多模态和检索增强,确保一致行为的复杂性。

英文摘要

We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in context that differs from what was stored parametrically during training. We document asymmetric biases: models tend to prefer in-context information about entities which appear in text, but prefer parametric information about entities which appear in images. We relate this asymmetry to the late representational alignment across modalities, showing that the longer processing time associated with resolving visual entities prevents the suppression of the model's usual factual recall mechanism, thus resulting in more parametric answers. Chain-of-thought reasoning does not appear to resolve the gap, but increasing the amount of visual information in the context does show an effect. These results illustrate the complexity of ensuring consistent behavior as models become increasingly multimodal and retrieval-augmented.

CommentsAccepted to EMNLP Findings 2026. Code and dataset are available at https://github.com/aparaselli/slow-to-see-slow-to-suppress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑