发表机构
University of Electronic Science and Technology of China; University of Michigan, Ann Arbor; Nanyang Technological University; Peking University; Mohamed bin Zayed University of Artificial Intelligence; New York University; Harvard University; Massachusetts Institute of Technology(电子科技大学; 密歇根大学安娜堡分校; 南洋理工大学; 北京大学; 穆罕默德·本·扎耶德人工智能大学; 纽约大学; 哈佛大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReMAP提出两种互补的潜在记忆(全局与局部)及强化学习访问策略,在推理时按需检索视觉证据,显著提升多图像基准性能并大幅减少视觉标记。
AI 中文摘要
随着多模态大语言模型(MLLMs)进行更长时间的推理,对初始视觉输入的注意力会减弱,从而削弱视觉锚定。视觉记忆在推理过程中重新引入视觉证据。我们从三个维度对视觉记忆进行了受控分析:策展、组织和访问。我们发现,局部证据受益于全局上下文,紧凑的潜在表示在准确性和视觉上下文成本之间取得了平衡,并且记忆访问的效用取决于推理状态。基于这些发现,我们提出了ReMAP(推理时记忆增强感知),它耦合了两种互补的潜在记忆:一个静态的、由问题条件化的全局记忆,用于保留场景和跨图像上下文;以及一个动态的局部记忆,它利用该上下文作为锚点,同时根据当前推理状态选择和重新编码区域级证据。两种记忆都返回插入到推理序列中的紧凑潜在标记,并且一个使用分支轨迹训练的强化学习访问策略决定何时继续推理或调用全局或局部记忆。在十个基准家族上,ReMAP在所有四个多图像基准上优于先前的视觉记忆方法,在MuirBench和MIMIC上分别超过最强先前结果8.38和14.84个百分点。在四个骨干家族中,启用记忆访问优于同一训练模型禁用记忆的情况,并且在共享的V*Bench、CV-Bench-2D和MuirBench问题上,ReMAP相对于原生分辨率骨干将进入推理序列的视觉标记减少了51.0-76.8%。进一步的分析表明,全局和局部记忆形成了不同但互补的潜在表示。这些组件共同通过让推理状态触发有针对性的视觉检索来恢复感知循环,检索到的证据指导后续推理。
英文摘要
As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that local evidence benefits from global context, compact latent representations balance accuracy and visual-context cost, and the utility of memory access depends on the reasoning state. Guided by these findings, we propose ReMAP (Reasoning-Time Memory-Augmented Perception), which couples two complementary latent memories: a static, question-conditioned Global memory that preserves scene and cross-image context, and a dynamic Local memory that uses this context as an anchor while selecting and re-encoding region-level evidence according to the current reasoning state. Both memories return compact latent tokens inserted into the reasoning sequence, and a reinforcement-learning access policy trained with branched rollouts decides when to continue reasoning or invoke Global or Local memory. On ten benchmark families, ReMAP outperforms prior visual-memory methods on all four multi-image benchmarks, exceeding the strongest prior results on MuirBench and MIMIC by 8.38 and 14.84 percentage points. Across four backbone families, enabling memory access improves over the same trained model with memory disabled, and on shared V*Bench, CV-Bench-2D, and MuirBench questions ReMAP reduces the visual tokens entering the reasoning sequence by 51.0-76.8% relative to the native-resolution backbone. Further analyses show that Global and Local memory form distinct yet complementary latent representations. Together, these components restore the perceptual cycle by letting the reasoning state trigger targeted visual retrieval, with the retrieved evidence guiding subsequent reasoning.
Comments32 pages. Code coming soon