arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Remember-R1:通过强化学习缓解长上下文视觉遗忘

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

Jianmin Chen, Jiaqi Tang, Wei Wei, Xiaogang Xu, Jiafei Wu, Zhe Liu, Qianzhou Wang, Yingying Yan, Botong Geng, Yuyang Xia, Lei Zhang, Qifeng Chen

arXiv 2608.01314首次发表:更新:

发表机构

Northwestern Polytechnical University; Hong Kong University of Science and Technology; Zhejiang University(西北工业大学; 香港科技大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态大语言模型的长上下文视觉遗忘问题,提出强化学习框架Remember-R1,通过过程级监督优化视觉依赖与注意力,提升了多模态推理性能。

AI 中文摘要

多模态大语言模型(MLLMs)越来越依赖长思维链推理完成复杂任务。然而,随着推理序列变长,模型可能逐渐减少对视觉证据的依赖,更多依赖累积的文本上下文,导致视觉遗忘。现有方法未直接约束视觉证据在原始推理轨迹中的使用与维护,长上下文视觉遗忘问题未得到充分解决。为解决该问题,本文提出Remember-R1,这是一个强化学习框架,通过直接在原始推理轨迹上应用过程级监督来缓解长上下文视觉遗忘。具体而言,Remember-R1引入奖励机制,鼓励更广泛覆盖匹配的视觉关键词、在后续推理步骤中更强地保持视觉依赖,以及更聚焦于问题相关的图像区域。在多个模型规模和不同多模态基准上开展的实验表明,Remember-R1可持续提升推理性能。额外分析进一步显示,它能减缓生成过程中视觉注意力的下降,证明其缓解长上下文视觉遗忘的有效性。

英文摘要

Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely less on visual evidence and more on accumulated textual context, leading to visual forgetting. Existing approaches do not directly constrain how visual evidence is used and maintained along the original reasoning trajectory, leaving long-context visual forgetting insufficiently addressed. To address this issue, we propose Remember-R1, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory. Specifically, Remember-R1 introduces rewards that encourage broader coverage of matched visual keywords, stronger persistence of visual dependence in later reasoning steps, and greater focus on question-relevant image regions. Experiments across multiple model scales and diverse multimodal benchmarks demonstrate that Remember-R1 consistently improves reasoning performance. Additional analyses further show that it slows the decline of visual attention during generation, supporting its effectiveness in mitigating long-context visual forgetting.

CommentsAccepted by ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑