AI 中文总结
该研究针对现有多模态长期智能体检索机制不足的问题,提出RRM框架,通过反思式经验记忆提炼跨任务检索策略,在多个多模态推理数据集上取得优于现有方法的性能。
AI 中文摘要
现有多模态长期智能体利用外部记忆克服长视频上下文有限的问题,但多数方法侧重存储内容而非存储记忆的检索方式。当检索不准确或多次无法获取有用证据时,现有智能体缺乏从过往任务轨迹诊断失败并调整未来搜索的机制。本文提出反思式检索记忆(Reflective Retrieval Memory,RRM),这是一种用于长周期多模态推理的反思式记忆框架。RRM在以实体为中心的多模态记忆图基础上,新增反思式经验记忆,该记忆从历史任务轨迹中提炼可迁移的程序性检索知识。与保存当前视频事实证据的情景记忆和语义记忆不同,反思式经验记忆捕获跨任务的可复用搜索策略。RRM将检索到的经验转化为查询级指导,而答案生成仅以从当前视频中新检索到的事实证据为条件。生命周期管理机制进一步通过使用频率、复用反馈和时间衰减调控经验记忆,从而减少冗余和噪声。RRM在M3-Bench-Robot、M3-Bench-Web和Video-MME-Long数据集上始终优于以往的最先进方法,证明了反思式检索记忆对长周期多模态推理的有效性。
英文摘要
Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When retrieval becomes inaccurate or repeatedly fails to obtain useful evidence, existing agents lack mechanisms to diagnose failures from previous task trajectories and adapt future search strategies.We introduce Reflective Retrieval Memory (RRM), a reflective memory framework for long-horizon multimodal reasoning. RRM augments an entity-centric multimodal memory graph with reflective experience memory, which distills transferable procedural retrieval knowledge from historical task trajectories. Unlike episodic and semantic memories that preserve factual evidence from the current video, reflective experience memory captures reusable search strategies across tasks. RRM converts retrieved experiences into query-level guidance, while answer generation remains conditioned only on factual evidence newly retrieved from the current video. A lifecycle management mechanism further regulates experience memory through usage frequency, reuse feedback, and temporal decay, thereby reducing redundancy and noise. RRM consistently outperforms previous state-of-the-art approaches on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, demonstrating the effectiveness of reflective retrieval memory for long-horizon multimodal reasoning.