发表机构
Wangxuan Institute of Computer Technology, Peking University; MemoraX AI(北京大学王选计算机研究所; MemoraX人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GraphMemix作为组合优化图记忆框架,通过三个关键组件构建查询感知证据森林,在四个多模态记忆基准上提升了准确率并建立了新的帕累托前沿。
AI 中文摘要
组织多模态智能体的长期记忆仍具挑战性,因为现有方法要么存在与查询无关的昂贵离线摘要问题,要么采用朴素的嵌入相似度匹配,会引入不完整且冗余的上下文。为解决这些问题,我们提出GraphMemix,这是一种组合优化图记忆框架,将记忆组织建模为查询感知证据森林的构建。具体而言,我们的方法包含三个关键组件:(1)候选图构建,通过模式和语义关系扩展多视图种子记忆,以获取查询感知的原始上下文;(2)证据效用与激活成本,将直接记忆支持与锚定条件关系验证解耦,从而抑制冗余或冲突信息;(3)森林优化,在最大证据预算及其可靠关系结构下,联合选择森林格式的记忆上下文。通过将记忆组织为与查询相关的子图,该方法避免了大量生命周期成本,并恢复了低相似度的互补证据。在四个长期多模态记忆基准上的实验结果表明,该方法在不同基础模型下均取得了显著改进,并在准确率与生命周期成本之间建立了新的帕累托前沿。
英文摘要
Organizing long-term memory for multimodal agents remains challenging because existing methods either suffer from expensive question-agnostic offline summaries or naive embedding similarity matching that introduces incomplete and redundant context. To address these issues, we propose GraphMemix, a combinatorial-optimization graph memory framework that models memory organization as query-aware evidence-forest construction. Specifically, our method consists of three key components:(1) candidate graph construction, which expands multi-view seed memories through schema and semantic relations to acquire query-aware original context; (2) evidence utility and activation costs, which decouples direct memory support from anchor-conditioned relation verification to suppress redundant or conflicting information; and (3) forest optimization, which jointly selects a forest-format memory context under a maximum evidence budget and its reliable relational structure. By organizing memory into a query-relevant subgraph, the method avoids substantial lifecycle cost and recovers low-similarity complementary evidence. Experimental results across four long-term multimodal memory benchmarks demonstrate significant improvements with different foundation models and establish a new Pareto frontier between accuracy and lifecycle cost.
CommentsProject page with code: https://github.com/ligeng0197/graphmemix