arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为视觉-语言-动作模型重构与优化潜在推理流

Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models

Hongyu Shi, Sen Zhao, Zuyu Zhang, Lifeng Shen, Ding Zou, Xinyu He, Xu Zhang, Qinghua Zhang

arXiv 2610.12090首次发表:更新:

发表机构

Chongqing University of Posts and Telecommunications; Academy of Advanced Interdisciplinary Studies; School of Computer Science and Technology; School of Artificial Intelligence; ZTE Corporation; Towngas(重庆邮电大学; 先进跨学科研究院; 计算机科学与技术学院; 人工智能学院; 中兴通讯股份有限公司; Towngas(香港中华煤气有限公司))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出FLOWMEM模型,通过动态检索重构潜在片段形成推理路径并优化,在RoboMME和LIBERO-Plus上分别提升1.7、4.1个百分点,验证了复用成功潜在计算对闭环VLA控制的价值。

AI 中文摘要

潜在推理使视觉-语言-动作(VLA)模型能够在生成连续机器人动作前,将多模态观测转化为与任务相关的内部状态。现有方法针对每个策略查询学习生成或优化此类状态,但执行后会丢弃成功的推理,因此需从头开始重建类似计算。本文提出Reasoning and Flow Memory(FLOWMEM),这是一种统一的VLA模型,可将成功的潜在计算转化为可复用的推理经验。FLOWMEM并非追加固定的检索上下文,而是随具身环境的变化动态检索并重构兼容的潜在片段,形成遵循成功计算时间结构与进展的推理路径。该路径随后会用当前视觉和本体感受证据进行优化,再用于条件化动作生成。在RoboMME和LIBERO-Plus上的实验显示,FLOWMEM的成功率分别达48.0%和77.3%,分别比无记忆策略高出1.7和4.1个百分点。这些结果证明了在闭环VLA控制中复用成功潜在计算的价值。

英文摘要

Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation. The route is then refined using current visual and proprioceptive evidence before it conditions action generation. Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively. These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑