DSV-Mem:评估MLLM智能体在专业工作流中的多模态记忆
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
- Amazon AGI(亚马逊AGI)
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对专业工作流中MLLM智能体多模态记忆评估的空白,提出DSV-Mem基准,含1000个问题与生成工具,发现状态演化是主要难点,最强基线得分低于45%,状态感知设计更有效。
中文摘要 AI 辅助
对话式多模态大语言模型(MLLM)智能体正日益被期望在专业工作流中提供协助,涵盖从AI研究与工程设计到产品管理与业务运营等领域。然而,这一能力仍未得到充分探索:现有基准大多聚焦于非正式、日常交互及个人生活场景,这些场景涉及摄影自然图像、孤立的静态工件以及基于回忆的问题。相比之下,专业场景通常涉及结构化、信息密集的工件,这些工件会经历频繁修订和权威更新,以及需要调和多个工件版本并精确跟踪状态的组合查询。为应对这些挑战,我们引入了DSV-Mem,一个用于评估密集状态视觉记忆(Dense Stateful Visual Memory)的基准。DSV-Mem包含专家审阅的场景和1000个问题,涵盖五个面向用户的类别(当前状态、过去状态、派生状态、变更历史和冲突/拒绝)。一个受Hartley启发的标准偏好于需要更广泛视觉证据检查需求的问题。我们还引入了一个生成工具,通过将状态转换合成与对话填充解耦来生成评估套件。对涵盖前沿和开放权重模型及记忆管理方法的27种配置的评估显示,最强基线在DSV-Mem上得分低于45%。分析揭示了以下发现:1)多模态性和信息密度都增加了难度,但状态演化,特别是控制更新的数量,是主要的测试因素。原始对话/草垛长度、OCR和算术不是主要瓶颈;2)模型在回答前常常未能根据先前的状态更新验证用户前提;3)增加推理努力和记忆管理方法带来的收益有限,而状态感知设计被证明更为有效。该基准和代码将公开发布。
英文摘要
Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.