面向多模态智能体学习者的双粒度智能体记忆与Shapley上下文归因
Dual-Grained Agent Memory and Shapley Context Attribution for Multimodal Agentic Learner
- UC Merced(加州大学默塞德分校)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态大模型推理不足的问题,提出双粒度智能体记忆框架DG-Mem,结合Shapley上下文归因,在多个数学多模态推理数据集上实现性能提升,且无需参数更新即可适配各类骨干模型。
AI中文摘要:
前沿多模态大语言模型(MLLM)具备出色的感知能力,但在科学与数学推理任务中仍表现不佳。对于闭源权重或设备端骨干模型,无法进行参数级适配,而无状态提示则会丧失已解决问题带来的复合收益。我们提出DG-Mem,一种双粒度智能体记忆框架,通过从训练时的rollout构建一次、测试时只读调用的非参数化外部存储记忆,对冻结的MLLM进行增强。受人类记忆的互补学习系统(CLS)理论启发,DG-Mem将其存储划分为基于实例的示例记忆和基于类别的IF-THEN规则的模式记忆,同时设置临时反思存储来协调二者的构建,确保模式仅由抽象反思合成,而非来自示例文本。DG-Mem有两项关键设计:一是在线概念分类器,在训练期间逐步扩展类别空间,而非采用预定义分类体系;二是Shapley上下文归因过程,将整个检索规则集的正确性分解,生成每条规则的效用值,用于测试时重新加权检索。该流程无需梯度更新,可部署在闭源权重或设备端骨干模型上。在MathVista、MMMU和MMMU-Pro数据集上,针对Qwen3.5-27B、Qwen3.5-122B-A10B、GPT-5-Nano、Gemini-3-Flash四种开源权重和专有骨干模型,DG-Mem相较于无记忆和竞争记忆基线均实现了持续提升。
英文摘要:
Frontier multimodal large language models (MLLMs) deliver impressive perception yet still falter on scientific and mathematical reasoning. Parameter-level adaptation is unavailable for closed-weight or on-device backbones, and stateless prompting forfeits any compounding benefit from problems already solved. We propose \textbf{DG-Mem}, a dual-grained agentic memory framework that augments a frozen MLLM with a non-parametric, externally stored memory built once from training-time rollouts and consulted read-only at test time. Motivated by the Complementary Learning Systems (CLS) account of human memory, DG-Mem factors its store into an instance-grounded exemplar memory and a category-level schema memory of IF-THEN rules, with a transient reflection store mediating their construction so that schemas are synthesized only from abstract reflections, never from exemplar text. Two design choices distinguish DG-Mem: an online concept categorizer that grows the category space incrementally during training rather than committing to a predefined taxonomy, and a Shapley context attribution procedure that decomposes correctness across the entire retrieved rule set and yields a per-rule utility that re-weights retrieval at test time. The pipeline introduces no gradient updates and is deployable on closed-weight or on-device backbones. Across MathVista, MMMU, and MMMU-Pro on four open-weight and proprietary backbones (Qwen3.5-27B, Qwen3.5-122B-A10B, GPT-5-Nano, Gemini-3-Flash), DG-Mem improves consistently over no-memory and competitive memory baselines.