发表机构
HKUST; ByteDance(香港科技大学; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出支持检索与分析记忆的AdaMM框架,在MemEye、MemGallery基准上分别提升多模态智能体长期记忆性能11.3%、7.3%,拓展了多模态记忆能力。
AI 中文摘要
长期多模态记忆不仅需要支持检索相关信息,还需支持对交互过程中积累的观测结果进行计算。现有系统大多强调检索记忆,通过摘要和索引组织交互历史,以返回不同粒度(从高层抽象到底层记录)的查询相关信息。本文将分析记忆(analytic memory)定义为一种互补抽象,它将重复出现的多模态观测结果组织成可查询结构,支持过滤、聚合、排序和时间比较。我们提出AdaMM框架,该框架同时支持检索记忆和分析记忆。AdaMM不依赖应用定义的模式,而是从对话、图像和上下文元数据中提取带有来源关联的属性-值观测结果,发现重复出现的字段结构并将其具体化以用于分析访问。推理时,感知记忆的规划器将查询分解为检索和分析操作,并将每个操作路由到适当的工具。在两个长期多模态记忆基准MemEye和MemGallery上进行的实验表明,AdaMM的性能分别提升了11.3%和7.3%。
英文摘要
Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking, and temporal comparison. We present AdaMM, a framework that jointly supports retrieval and analytic memory. Rather than relying on application-defined schemas, AdaMM extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access. At inference time, a memory-aware planner decomposes queries into retrieval and analytic operations and routes each operation to the appropriate tools. Experiments on two long-term multimodal memory benchmarks, MemEye and MemGallery, show that AdaMM improves performance by up to 11.3\% and 6.9\%, respectively.