AI 中文总结
研究多模态智能体记忆投递阶段,发现检索与回答间存在未隔离环节,提出DeliverMem方法,通过保留原始模态、赋予身份标识和注明时间三个决策提升性能,在MemLens和DMV-Bench基准上均领先现有方法。
AI 中文摘要
关于多模态智能体记忆的研究优化了写入、更新和检索的内容。然而,在检索与回答之间,存在一个多模态记忆评估未单独隔离的阶段:检索到的记忆中有多少到达模型,以及以何种形式到达。我们称之为“投递”,在MemLens上的受控分解实验定位了剩余提升空间所在。在检索到的证据集完全固定的情况下,投递原始像素而非扣留它们,在8B骨干模型上准确率提升了13.87个百分点,而仅优化同一批消息的检索完美性仅提升2.31个百分点。在所有三个MemLens骨干模型上,投递是更大的贡献项,且随骨干模型能力增强而增长;检索的贡献也在增长,但未能缩小差距。我们提出DeliverMem,将投递实例化为三个决策:保留原始模态、为每个条目赋予可读的身份标识、并注明其被看到的时间,同时为投递无法提供的属性配备一个检索端适配器。每个决策均与仅改变自身变量的投递匹配对照进行测量。DeliverMem在所有四种上下文长度下均领先MemLens上最强的已发表记忆智能体,并在两个骨干模型的所有设置下均优于DMV-Bench自身的最强方法。在MemLens上,它仅使用十分之一到七十分之一的输入量即可实现这一结果。每个决策仅在问题缺少其所提供信息时发挥作用,在其他情况下则无效。尽管如此,一个固定的配置在两个基准上均领先,无需训练任何组件或修改存储记录。项目页面:此https URL
英文摘要
Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery, and a controlled decomposition on MemLens locates the remaining room there. With the retrieved evidence set exactly fixed, delivering the original pixels instead of withholding them raises accuracy by 13.87 points on an 8B backbone, whereas making retrieval perfect on those same messages improves it by 2.31. Delivery is the larger term on all three MemLens backbones and grows with backbone strength; retrieval grows too, without closing the gap. We propose DeliverMem, an instantiation of delivery as three decisions: keep the original modality, give each item a readable identity, and state when it was seen, with a retrieval-side adapter for the one property delivery cannot supply. Each is measured against a delivery-matched control that alters only its own variable. DeliverMem leads the strongest published memory agent on MemLens at all four context lengths, and beats DMV-Bench's own strongest method at every setting on both backbones. On MemLens it does this on a tenth to a seventieth of the input. Each decision helps only where the question lacks what it supplies, and is null elsewhere. A single fixed configuration nonetheless leads both benchmarks, without training any component or modifying the stored records. Project page: https://avalon-s.github.io/DeliverMem/
Comments32 pages, 6 figures, 21 tables. Project page: https://avalon-s.github.io/DeliverMem/