arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37307cs.RO

记住你所做的:用于长视界视觉-语言-动作策略的动作历史记忆与双专家去噪

Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

发表机构哈尔滨工业大学 · 香港大学 · 香港理工大学
查看机构详情
  • Harbin Institute of Technology(哈尔滨工业大学)
  • The University of Hong Kong(香港大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yaxin Zhao, Dianye Huang, Chenwei Wang, Chenguang Yang, Zhongliang Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

提出ActMem-VLA,通过Mamba记忆模块和预动作专家双专家架构,为冻结的VLA补充动作历史,解决感知混淆,在LIBERO-Mem上以3.45%额外参数提升成功率至80.8%。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型推动了机器人操作领域的快速发展,展现出强大的细粒度控制能力,并在长视界任务上取得了有前景的性能。然而,许多现有的VLA缺乏对交互历史的显式访问,使其容易受到感知混淆的影响:在不同任务阶段,相似的当前观察和机器人状态可能引发动作歧义并降低成功率。现有方法通过特征条件化、动作先验修改或采样引导来整合时间或进度线索。然而,联合微调记忆模块和基础VLA的方法会增加额外的策略训练成本,这促使我们将可训练的历史条件化引导与冻结的基础策略细化分离。我们提出ActMem-VLA,一种双专家交接架构,通过一个记忆插件增强冻结的、微调后的VLA,该插件包含一个基于Mamba的记忆模块和一个轻量级预动作专家(PAE)。具体而言,Mamba将已执行动作的历史编码为记忆,该记忆与当前上下文一起条件化PAE。利用这些输入,PAE在早期高噪声去噪阶段引导任务进展,然后将部分去噪的动作传递给冻结的动作专家(AE),在剩余低噪声步骤中细化动作细节。在整个训练过程中,微调后的基础VLA保持冻结,仅联合优化Mamba模块和PAE。在LIBERO-Mem上,ActMem-VLA在全部十个任务上实现了80.8%的平均成功率,而π0.5为65.2%,MemoryVLA为49.5%,同时仅引入了3.45%的额外参数。在四个真实世界任务中,它比π0.5的平均成功率提高了28.8%。

英文摘要

Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8\% average success across all ten tasks, compared with 65.2\% for $π_{0.5}$ and 49.5\% for MemoryVLA, while introducing only 3.45\% additional parameters. Across four real-world tasks, it improves the average success rate over $π_{0.5}$ by 28.8\%.

↑