arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18704cs.CLcs.AI

MemFuse:基于碎片化观测的多源记忆融合

MemFuse: Multi-Source Memory Fusion from Fragmented Observations

Chao Li, Yuanfa Li, Wenhao Wu, Xule Liu, Zhi Wang, Kun Shao

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有记忆系统多聚焦单源文本、无法处理碎片化信息的问题,提出多源记忆融合基准MemFuseBench与结构化记忆系统MemFuse,实验显示MemFuse在跨源证据融合任务上表现最优。

中文摘要 AI 辅助

长期记忆对于在长时间交互中运作的智能体至关重要,但现有的记忆系统和基准测试大多聚焦于单源文本历史。然而在现实场景中,相关信息往往分散在不同应用、设备、用户及时间中,要求智能体将分散的观测整合为连贯的情景记忆,同时保留其来源出处。为解决这些不足,我们推出多源记忆融合基准测试MemFuseBench。MemFuseBench采用场景到传感器的流水线,将可控场景合成为带来源标签的观测、有证据支撑的问题以及对抗性干扰项,可用于系统评估时间推理、跨源证据融合以及抗噪声能力。我们进一步提出结构化记忆系统MemFuse,其在事件层原子记忆中保留源级证据,并在因果融合图中将相关原子事件组织为簇层融合记忆;检索时,MemFuse会检索并组织相关证据片段,同时保持对原始源事件的可追溯性。在MemFuseBench上的实验表明,MemFuse在所有三种大语言模型(LLM)设置下的被评估记忆系统中取得了最佳整体性能,且在需要跨源证据融合的问题上持续提升表现。

英文摘要

Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantly focus on single-source textual histories. In realistic settings, however, relevant information is often fragmented across applications and devices, as well as across users and time, requiring agents to integrate dispersed observations into coherent episodic memories while preserving their source provenance. To address these gaps, we introduce **MemFuseBench**, a benchmark for *multi-source memory fusion*. MemFuseBench is built with a Scene-to-Sensor pipeline that synthesizes controllable scenarios into source-tagged observations, evidence-grounded questions, and adversarial distractors. It enables systematic evaluation of temporal reasoning, cross-source evidence fusion, and robustness to noise. We further propose **MemFuse**, a structured memory system that preserves source-level evidence in event-layer atomic memory and organizes related atomic events into cluster-layer fused memory within a causal fusion graph. During retrieval, MemFuse retrieves and organizes related evidence fragments while maintaining traceability to original source events. Experiments on MemFuseBench show that MemFuse achieves the best overall performance among the evaluated memory systems under all three LLM settings and consistently improves performance on questions requiring cross-source evidence fusion.

补充信息

↑