看过、说过还是被遗忘了?跨对话轮次的视觉键值记忆因果审计
Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns
浏览论文内容
中文总结 AI 辅助
研究跨对话轮次视觉键值记忆何时可安全遗忘,提出因果视觉记忆审计框架,通过在VisDial和ConvBench上实验发现当前注意力对未来有用区域排名不佳,揭示安全遗忘受低未来视觉依赖性或特定事实语言表达支持而非低当前注意力。
中文摘要 AI 辅助
有状态的多模态助手对图像进行一次编码,但可能在许多轮次后才回答关于它的问题。注意力引导的视觉键值逐出假设现在无关的证据将来仍无用,尽管未来的问题未知。我们探究视觉事实何时真正可以安全遗忘,并引入因果视觉记忆审计(CVMA),这是一个配对单预填充框架,测试当一个视觉区域、整个图像或先前的助手文本不可用时,后续答案会失去什么。在VisDial和ConvBench上,即使诊断边际效用控制显示有很大选择空间,当前注意力对未来有用区域的排名可能比随机还差。当后续轮次不需要视觉时,总体分数会掩盖这种失败;受控和随机生成的历史揭示了另一条出路,即对于已陈述的事实,助手文本键值会取代图像键值,但对于未陈述的事实则不可靠。在测试的堆栈中,安全遗忘由低未来视觉依赖性或特定事实的语言表达支持,而非低当前注意力。
英文摘要
Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes that evidence irrelevant now will remain dispensable, although future questions are unknown. We ask when a visual fact is actually safe to forget and introduce the Causal Visual Memory Audit (CVMA), a paired single-prefill framework that tests what later answers lose when a visual region, the whole image, or prior assistant text becomes unavailable. On VisDial and ConvBench, current attention can rank future-useful regions worse than random even though a diagnostic marginal-utility control shows substantial selection headroom. Aggregate scores hide this failure when later turns do not need vision; controlled and stock-generated histories reveal a second escape route, in which assistant-text KV replaces image KV for facts already stated but not reliably for unstated facts. In the tested stacks, safe forgetting is supported by low future visual dependence or fact-specific verbalization---not by low current attention.