UniMem:统一视觉-语言-动作模型的多模态记忆与控制
UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
UniMem 是统一多模态记忆与控制的视觉-语言-动作模型框架,通过事件分类器、关键帧编码器等技术,在模拟和硬件任务中性能优于基线,推理更快、训练流程更简单。
中文摘要 AI 辅助
尽管视觉-语言-动作(VLA)模型已利用互联网规模的预训练和面向任务的微调在长 horizon 任务上取得了优异性能,但它们通常在需要记忆的非马尔可夫任务上表现不佳。现有的记忆方法通常涉及额外的视觉-语言模型(VLM)用于长期记忆管理,这会引入记忆瓶颈和碎片化的训练流程。基于多个历史帧进行条件设置可为 VLA 模型提供过去场景更具描述性的特征,但如果以任意固定间隔选择帧,则会降低性能。为解决这些局限,我们提出了 UniMem,一个在单一主干下统一高级多模态记忆和低级控制的框架。UniMem 采用事件分类器进行记忆更新、关键帧编码器用于密集空间记忆,以及关键帧缓存技术以最小化策略 rollout 期间的开销。我们在 5 项模拟任务和 4 项硬件任务(针对顺序和空间记忆)上评估 UniMem,结果表明,我们的统一单模型系统在模拟任务中优于固定间隔图像采样基线(93.4% 对 68.2%),在硬件任务中优于分层基线(80.0% 对 43.5%),同时提供更快的推理速度和简单的训练流程以方便采用。项目网站:this https URL
英文摘要
While Vision-Language-Action (VLA) models have leveraged internet-scale pretraining and task-focused finetuning to achieve strong performance on long-horizon tasks, they often struggle with non-Markovian tasks that require memory. Existing approaches to memory typically involve additional Vision-Language-Models (VLMs) for long-term memory management, introducing a memory bottleneck and a fractured training pipeline. Conditioning on multiple historical frames can provide the VLA with access to more descriptive features of past scenes, but can degrade performance if frames are chosen at arbitrary, fixed intervals. To address these limitations, we present UniMem, a framework that unifies high-level, multimodal memory and low-level control under one backbone. UniMem employs an event classifier for memory updates, a keyframe encoder for dense spatial memory, and a keyframe caching technique to minimize overhead during policy rollouts. We evaluate UniMem across five simulation and four hardware tasks targeting sequential and spatial memory, demonstrating that our unified, single-model system outperforms fixed-interval image sampling baselines (93.4% vs. 68.2%) in simulation and hierarchical baselines (80.0% vs. 43.5%) in hardware, while offering faster inference and a simple training pipeline for easy adoption. Project website: https://losterberg3.github.io/unimem-vla/
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。