arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01456cs.CVcs.CL

基于多模态记忆压缩的长视野具身决策

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression

Bingxuan Li, Rui Yang, Cheng Qian, Jiateng Liu, Jeonghwan Kim, Zhenhailong Wang, Manling Li, Tong Zhang, Heng Ji

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出DunphyBench基准用于评估长视野具身决策,发现当前智能体与人类表现存在差距,设计了MeMento记忆压缩器,使VLM驱动智能体准确率提升7.18%且内存使用降低85.38%。

中文摘要 AI 辅助

智能体日益被期望不仅作为任务执行者,还作为代表人类用户的决策者。这一转变要求智能体在长视野中积累证据、解读隐性用户偏好,并在部分观测下比较多个候选对象。本研究提出DunphyBench,这是一个用于评估智能体在以人类为中心的长视野具身决策任务的新基准,智能体需在多个具身化住房环境中导航,并做出符合多维度人类偏好的决策。与通常关注程序规划或即时目标完成的标准具身推理任务不同,该设置要求智能体将多模态、多源输入整合为连贯知识,以支持长视野下的复杂推理。评估结果显示,当前智能体与人类表现之间存在显著差距。此外,我们对最先进的视觉语言模型(VLM)驱动智能体的诊断表明,记忆管理是瓶颈之一,原始多模态历史会引入噪声,损害决策质量。基于这一发现,我们设计了MeMento,这是一种偏好条件下的多模态记忆压缩器,它基于用户偏好,利用固定数量的记忆标记从长视野历史中选择性压缩与决策相关的信息。实验表明,与最强基线相比,MeMento使VLM驱动智能体的准确率提升7.18%,同时内存使用量降低85.38%。

英文摘要

Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents to accumulate evidence over long horizons, interpret implicit user preferences, and compare multiple candidates under partial observations. In this work, we propose DunphyBench, a new benchmark for evaluating agents on long-horizon human-centered embodied decision-making, where the agent must navigate through multiple embodied housing environments and make decisions that align with multi-dimensional human preferences. Unlike standard embodied reasoning tasks that often focus on procedural planning or immediate goal completion, our setting requires agents to integrate multimodal, multi-source input into coherent knowledge that supports complex reasoning across long horizon. The evaluation results reveal that there is a substantial gap between current agents and human performance. Furthermore, our diagnosis of state-of-the-art VLM-driven agents reveals that memory management is one of the bottlenecks, where raw multimodal history introduces noise that hinders decision quality. Motivated by this finding, we design MeMento, a preference-conditioned multimodal memory compressor that selectively compresses decision-relevant information from long-horizon history based on user preferences with a fixed set of memory tokens. Experiments show that MeMento helps VLM-driven agents improve accuracy by 7.18%, while reducing memory usage by 85.38% compared to the strongest baseline.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑