V-Mem:面向多模态智能体长期记忆的模态路由检索
V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
浏览论文内容
中文总结 AI 辅助
该研究针对多模态智能体记忆的模态与相似性-相关性差距,提出V-Mem系统,通过模态路由检索优化,在Mem-Gallery、LoCoMo数据集上显著提升多模态问答性能。
中文摘要 AI 辅助
用户与大语言模型(LLM)智能体的交互日益呈现多模态特征:对话中文字与图像交错出现,后续问题可能针对任一模态内容。然而,大多数智能体记忆系统围绕文本设计,即便少数支持多模态对话的系统,在视觉相关问题上仍表现不佳。我们将这种失败归因于其依赖的相似性搜索背后的假设:在索引空间中,查询与能回答它的相关证据距离较近。在多模态场景下,两个差距会打破这一假设。模态差距:即使在经过训练的联合嵌入空间中,查询与自身模态的记忆内容的距离也会比与另一模态证据的距离更近。相似性-相关性差距:与查询最相似的内容往往并非能回答该查询的证据,当查询同时包含文字和图像、且其证据与单独任一模态部分都不相似时,这种情况最为明显。我们提出V-Mem,一种通过查询和目标证据的模态(均仅从查询中识别)来路由检索的多模态智能体记忆系统。为跨越模态差距,V-Mem将对话组织成多轮,返回与查询同轮的目标模态内容,无需跨模态比较。为缩小相似性-相关性差距,它使用LLM生成的锚点进行搜索,该锚点比查询更接近相关证据:对于寻找图像的纯文本查询,使用假设性描述作为锚点;当证据需结合文本与图像才能获取时,使用查询文本加上从查询图像中提取的相关关键词作为增强搜索锚点。在Mem-Gallery数据集上,V-Mem获得LLM评判分数0.82,而第二名仅为0.56,在带图像的问题上优势最大(0.87,无基准超过0.47);在LoCoMo数据集上,其分数为0.69,对比基准0.58。
英文摘要
Interaction between users and LLM agents is increasingly multimodal: conversations interleave text with images, and a later question may target either. Yet most agent memories are designed around text, and even the few that support multimodal conversations still fail on vision-related questions. We trace this failure to an assumption behind the similarity search they rely on: in the index space, a query lies close to the relevant evidence that answers it. In multimodal settings, two gaps break it. By the modality gap, a query lies closer to memory content of its own modality than to evidence in another, even in a trained joint embedding space. By the similarity-relevance gap, the content most similar to a query is often not the evidence that answers it, most acutely when a query carries both text and image and its evidence resembles neither part alone. We present V-Mem, a multimodal agentic memory system that routes retrieval by the modality of the query and that of the target evidence, both recognized from the query alone. To cross the modality gap, V-Mem organizes the conversation into rounds and returns the target-modality content from the same round as the match, without comparing across modalities. To close the similarity-relevance gap, it searches with an LLM-generated anchor that sits closer to the relevant evidence than the query does: a hypothetical caption for a text-only query seeking an image, and an enriched search anchor, the query text plus relevant keywords extracted from the query image, when the evidence is reachable only by combining the two. On Mem-Gallery, V-Mem reaches an LLM-judge score of 0.82 versus 0.56 for the second best, with the largest margin on questions carrying an image (0.87, no baseline above 0.47); on LoCoMo it scores 0.69 versus 0.58.