arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Divide-and-Remember:用于长时程VLA策略的递归动作相关记忆

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh

arXiv 2610.00982首次发表:更新:

发表机构

National University of Singapore; Nanyang Technological University, Singapore; LISTENAI(新加坡国立大学; 新加坡南洋理工大学; 聆心智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型在长时程操作任务中记忆不足的问题,提出Divide-and-Remember递归记忆方法,通过最大化条件互信息并采用共享选择器的递归top-K选择,在RoboMME基准上以64个token预算取得最先进成功率。

AI 中文摘要

视觉-语言-动作(VLA)模型在依赖历史信息的操作任务中表现不佳,这类任务中仅凭当前观测无法决定动作,策略需要记忆历史信息。现有记忆方法通过设计决定记忆内容,例如保留像素变化大的帧,但不同任务上的提升效果不一致。我们将记忆什么视为一个优化问题。从模仿学习的POMDP表述出发,我们证明最优记忆最大化在给定当前观测条件下动作与记忆之间的条件互信息$I(a_t; m_t \mid o_t)$。直观上,这意味着保留历史中与动作相关且当前观测未包含的信息。基于我们的分析,我们提出Divide-and-Remember(D&R),一种递归记忆方法,学习记忆函数$m_t = M(h_t)$,可扩展到长上下文同时保持计算轻量。它包含两种策略:(1)对完整历史的选择被递归划分为对$2K$个token进行top-$K$选择的子问题,从而端到端学习的固定大小轻量选择器支持无界历史;(2)所有递归块共享一个选择器,捕获每个块共有的选择规则并保持方法高效。在RoboMME(一个包含16个需要记忆何时、何地、何种及如何行动的长时程操作任务的基准)上,D&R在仅64个token的预算下取得了最先进的平均成功率,并在所有四个套件上获得一致提升;真实机器人实验显示相同增益。代码、检查点和更多结果见https URL。

英文摘要

Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information $I(a_t; m_t \mid o_t)$ between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function $m_t = M(h_t)$ and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-$K$ selection over $2K$ tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑