arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于视觉-语言-动作模型长程规划的显式语言记忆

Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models

Houze Xu, Jizhong Li, Ziyi Ye

arXiv 2608.04765首次发表:更新:

发表机构

Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言-动作模型长程规划的挑战,提出带显式语言记忆模块的分层架构,经仿真与实机实验验证,可提升复杂长程任务的成功率、鲁棒性与决策可解释性。

AI 中文摘要

视觉-语言-动作(VLA)模型提供了连接视觉感知、语言理解与机器人控制的统一范式。然而现有VLA模型在长程任务中仍面临重大挑战:稀疏的专家演示限制了跨任务的组合泛化能力;长程任务的非马尔可夫特性使得仅以当前观测为条件的策略难以维持时间一致性;有限的闭环纠错能力会导致执行误差累积;端到端的动作微调可能削弱视觉-语言模型(VLM)主干的高级语义表示。为解决这些问题,我们提出了一种带有显式语言记忆模块的分层长程VLA架构,核心思路是将离散的时间观测转换为带有时间逻辑的连贯文本记忆序列。该系统解耦为高级VLM与低级VLA:高级VLM通过视觉问答训练范式进行语义推理,低级VLA则在子任务指令与视觉观测的条件下执行精确的连续控制。高级VLM以先前的记忆为上下文锚点,递归更新语言记忆与子任务指令,从而在长程执行过程中实现持久的时间跟踪与动态修正。我们在多个仿真环境中对所提方法进行评估,并在真实机器人平台上开展了仿真到现实的实验。结果表明,显式语言记忆可提升VLA模型在复杂长程任务上的成功率与鲁棒性,同时为决策过程提供可解释的语义说明。

英文摘要

Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse expert demonstrations constrain cross-task compositional generalization; the non-Markovian nature of long-horizon tasks makes it difficult for policies conditioned only on current observations to maintain temporal consistency; limited closed-loop error correction allows execution errors to accumulate; and end-to-end action fine-tuning may weaken the high-level semantic representations of vision-language model (VLM) backbones. To address these issues, we propose a hierarchical long-horizon VLA architecture with an explicit language-memory module. The central idea is to convert discrete temporal observations into a coherent textual memory sequence with temporal logic. The system is decoupled into a high-level VLM and a low-level VLA: the high-level VLM performs semantic reasoning through a visual question answering training paradigm, while the low-level VLA executes precise continuous control conditioned on subtask instructions and visual observations. The high-level VLM recursively updates both language memory and subtask instructions using the previous memory as a contextual anchor, enabling persistent temporal tracking and dynamic correction during long-horizon execution. We evaluate the proposed method in multiple simulation environments and conduct sim-to-real experiments on a real robotic platform. The results demonstrate that explicit language memory improves the success rate and robustness of VLA models on complex long-horizon tasks while providing an interpretable semantic account of the decision process.

CommentsThis submission has been withdrawn by the authors, because the manuscript was uploaded to arXiv without the awareness of the remaining co-authors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑