arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CUE-Mem:通过多模态对话中的隐式线索对长期用户记忆进行基准测试

CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

Yulin Hu, Yanyan Zhao, Zimo Long, Xing Fu, Mengtong Ji, Weixiang Zhao, Yutai Hou, Qianchao Wang, Dandan Tu

arXiv 2609.32574首次发表:更新:

AI 中文总结

CUE-Mem是一个多模态基准,通过隐式线索评估长期用户记忆,包含2674个问题和四个任务,发现保留细微线索是主要瓶颈,并测试了文本化与原生多模态访问的效果。

AI 中文摘要

长期记忆对于与用户进行持续对话的多模态智能体至关重要。然而,用户记忆并不总是被明确陈述:它们也可能通过图像中反复出现的背景物体、音频中的环境声音或其他外围多模态线索被隐含地表达。现有基准大多聚焦于纯文本记忆或明确的多模态证据,使得隐式多模态线索的探索不足。我们引入了CUE-Mem,一个用于从隐式线索评估长期用户记忆的文本-图像-音频基准。CUE-Mem包含2,674个问题,覆盖明确和隐式证据设置,并涵盖四个任务:实体回忆、长模式、个性化推荐和答案拒绝。在文本化记忆系统中,隐式性能仍远低于oracle证据,主要瓶颈在于保留和检索细微线索,而非问题的可回答性。增加字幕细节能恢复更多此类证据,但带来不均衡的收益和迅速增长的token成本,这促使采用原生多模态访问。然而,原生访问并不能统一解决瓶颈:证据的使用强烈依赖于骨干网络,而多模态索引引入了大量检索噪声。CUE-Mem为选择性保留、检索和使用细微多模态证据的记忆系统提供了一个测试平台。

英文摘要

Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Existing benchmarks largely focus on text-only memory or explicit multimodal evidence, leaving implicit multimodal cues underexplored. We introduce CUE-Mem, a text-image-audio benchmark for evaluating long-term user memory from implicit cues. CUE-Mem contains 2,674 questions across explicit and implicit evidence settings and covers four tasks: Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal. Across textualized memory systems, implicit performance remains far below oracle evidence, locating the main bottleneck in preserving and retrieving subtle cues rather than question answerability. Increasing caption detail recovers more of this evidence, but brings uneven gains and rapidly growing token costs, motivating native multimodal access. Yet native access does not uniformly resolve the bottleneck: evidence use depends strongly on the backbone, while multimodal indexing introduces substantial retrieval noise. CUE-Mem provides a testbed for memory systems that selectively retain, retrieve, and use subtle multimodal evidence.

Comments28 pages. Submitted to AAAI 2027. Code: https://github.com/yulinlp/CUE-MEM. Data: https://huggingface.co/datasets/Kkryptonite/CUE-Mem

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑