arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SimpleMemVLA:一种简单但有效的视觉-语言-动作模型原生视频记忆

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Cheng Yin, Wang Xu, Junpeng Yang, Sikyuen Tam, Hanyu Liu, Yuan Yao, Xiangrui Zeng, Junbo Cui, Yequan Wang, Zhouping Yin, Yankai Lin

arXiv 2609.05533首次发表:更新:

发表机构

Huazhong University of Science and Technology; Zhongguancun Academy; Tsinghua University; Modelbest; Peking University; Renmin University of China; Beijing Academy of Artificial Intelligence(华中科技大学; 中关村学院; 清华大学; Modelbest; 北京大学; 中国人民大学; 北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SimpleMemVLA提出无需专用记忆模块的VLA,通过保留完整历史并以时间戳视频格式输入骨干网络,利用子任务隐藏状态传递信息,在四个记忆基准上达到最先进性能,且不损失通用控制能力。

AI 中文摘要

长时程操作是部分可观测的:选择下一个动作所需的信息可能只出现在几分钟前的观测中。现有的记忆机制:检索库、学习型压缩器、循环状态,必须在知道未来决策需要什么之前决定从过去保留什么。这是基于分钟级历史数据太大而无法直接处理的假设,而现代VLM骨干网络已不再如此。在这项工作中,我们引入了SimpleMemVLA,一种没有专用记忆模块的VLA。它保持采样的历史完整,并以骨干网络预训练处理的时间戳视频格式将其传递给骨干网络;生成子任务的隐藏状态随后形成从历史到标准流匹配动作头的唯一通道。由于连续决策共享大部分历史,在动作执行期间预填充共享前缀使延迟接近单帧VLA。SimpleMemVLA在四个记忆基准上创下了新的最先进水平,且不牺牲通用控制性能。在固定骨干网络和训练设置的情况下,它大幅优于检索、压缩和循环状态机制,因果干预证实策略确实读取了其历史。代码可在以下网址获取:https://this URL

英文摘要

Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, must decide what to keep from the past before knowing what a future decision will require. They were motivated by the assumption that minute-scale history is too large to process directly, which no longer holds for modern VLM backbones. We propose SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory. It keeps the sampled history intact in the timestamped video format the backbone was pretrained to process, routes the evidence it finds to a standard flow-matching action head through the hidden states of a generated sub-task, and prefills the history shared by consecutive decisions during action execution, keeping latency close to that of a single-frame VLA. SimpleMemVLA achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-state methods by a wide margin. History interventions show that the policy reads specific evidence from its past and follows edited histories without parameter updates, a visual form of in-context learning. On a physical dual-arm robot, it completes two tasks whose decisive evidence disappears before the robot acts. Code available at https://github.com/OpenBMB/SimpleMemVLA

Comments30 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑