AI 中文总结
针对现有LVLM智能体记忆系统的缺陷,提出PMMC框架,将部分记忆推理移至整合阶段,构建结构化问题库,提升答案质量与证据召回率,降低查询成本。
AI 中文摘要
长期记忆对于LVLM智能体在跨扩展多模态交互中保持一致性和整合信息至关重要。然而,现有的智能体记忆系统常将视觉体验简化为文本摘要,或依赖静态的“检索后推理”流程,这在查询时效率低下,且在问题需要图像-文本绑定、时间更新或视觉细节时表现脆弱。我们提出Prospective Multimodal Memory Compilation(前瞻性多模态记忆编译)框架,该框架将部分记忆推理过程从查询阶段转移至记忆整合阶段。给定累积的多模态交互,Questioner(提问器)预测未来的问题候选,Planner(规划器)编译基于问题条件的多模态记忆程序,Doubter(验证器)则验证规划的证据路径是否能支持预测的答案。经验证的问题-程序对构成结构化问题库,用于高效的查询时路由和证据检索。在多模态长期记忆基准上的实验表明,我们的方法提升了答案质量和视觉证据召回率,同时降低了查询时的token数和延迟成本。大量 ablation( ablation分析即消融实验)分析了自反馈、动态规划、原始图像访问及问题库覆盖率的影响。
英文摘要
Long-term memory is essential for LVLM agents to maintain consistency and integrate information across extended multimodal interactions. Existing agent memory systems, however, often reduce visual experiences into textual summaries or rely on static retrieve-then-reason pipelines, which are inefficient at query time and brittle when questions require image-text binding, temporal updates, or visual details. We propose Prospective Multimodal Memory Compilation, a framework that shifts part of the memory reasoning process from query time to memory consolidation time. Given accumulated multimodal interactions, a Questioner predicts future question candidates, a Planner compiles question-conditioned multimodal memory programs, and a Doubter verifies whether the planned evidence path can support the predicted answer. The verified question-program pairs form a structured question bank for efficient query-time routing and evidence retrieval. Experiments on multimodal long-term memory benchmarks show that our method improves answer quality and visual evidence recall while reducing query-time token and latency costs. Extensive ablations analyze the effects of self-feedback, dynamic planning, raw-image access, and question bank coverage.