发表机构
The University of Hong Kong; Southern University of Science and Technology(香港大学; 南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出D$^2$-VLA,通过双记忆与双频KV缓存机制,在长时动态操作中高效利用历史视觉线索,显著提升多任务成功率。
AI 中文摘要
长时程操作要求机器人在响应移动物体的同时,记住不再处于视野中的线索。然而,视觉-语言-动作(VLA)策略通常依赖最新的观测,刷新其视觉上下文通常需要另一次昂贵的视觉-语言模型(VLM)前向传播。我们提出D$^2$-VLA,它在预训练VLA的KV缓存接口处结合了双记忆和双频控制。D$^2$-VLA使用分块因果KV缓存来增量编码观测,并在不同时间注意力模式的引导下,为VLM和动作专家构建分离的历史KV读取视图。在周期性的VLM更新之间,一个门控适配器将新的视觉特征融入最新的历史条件KV块中,而一个短时快速记忆队列支持动作重规划。我们引入了DOMINO-Long,一个包含十个任务的基准,要求机器人在操作移动物体时使用早期视觉线索。D$^2$-VLA在DOMINO上实现了29.3%的完整任务成功率,相比之下,$\pi_{0.5}$为9.6%,PUMA为17.2%;在DOMINO-Long上达到60.0%,而两者分别为35.4%和20.6%。它在八个真实机器人任务上提高了成功率,并在LIBERO-Long上达到97.5%,在RoboTwin 2.0上达到74.3%。
英文摘要
Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D$^2$-VLA, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. D$^2$-VLA uses block-wise causal KV caching to encode observations incrementally and, guided by distinct temporal attention patterns, constructs separate historical KV read views for the VLM and action expert. Between periodic VLM updates, a gated adapter incorporates fresh visual features into the latest history-conditioned KV block, while a short fast-memory queue supports action replanning. We introduce DOMINO-Long, a ten-task benchmark requiring robots to use earlier visual cues when manipulating moving objects. D$^2$-VLA achieves complete-task success rates of 29.3\% on DOMINO, compared with 9.6\% for $π_{0.5}$ and 17.2\% for PUMA, and 60.0\% on DOMINO-Long, compared with 35.4\% and 20.6\%, respectively. It improves success rates on eight real-robot tasks and reaches 97.5\% on LIBERO-Long and 74.3\% on RoboTwin 2.0.
Comments30 pages