arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33412cs.CVcs.AI

解决状态表示不匹配:面向开环VLA规划的状态空间视觉推理

Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning

Junhao Xiao, Haoxiang Zhao, Menghao Fang, Jinkui Zhang, Jinghan Yu, Xinyu Huang, Zhiyu Wu, Kaiming Xu, Yi Chen, Youjun Bao, Zhiyuan Ma

首次发表
浏览论文内容

中文总结 AI 辅助

针对开环VLA规划中的状态表示不匹配问题,提出SSVR方法,通过解耦静态上下文与循环潜在状态,实现高效多步推理,显著提升规划性能。

中文摘要 AI 辅助

尽管视觉-语言-动作(VLA)模型取得了快速进展,现有推理范式在开环规划中仍面临根本性的“状态表示不匹配”问题。在仅给定初始观测的情况下,模型必须在内部模拟动作条件下的状态转移,而文本、像素和潜在空间推理分别可能遭受有损空间压缩、误差累积的视觉生成以及中间潜在标记的绕过,从而削弱了可靠的长时程规划能力。我们提出状态空间视觉推理(SSVR),该方法将静态视觉上下文、语言约束和循环潜在状态解耦。SSVR对初始图像和指令仅编码一次,然后基于潜在状态调节每个动作预测,并通过动作条件下的GRU更新该状态。以Qwen2.5-VL作为骨干模型,SSVR在FrozenLake、Maze和MiniBehavior上分别达到99.5/99.6、96.3/98.0和83.9/90.6的EM/PR,显著优于先前方法。大量实验支持了循环状态建模在VLA开环规划中跨输入变换和迁移设置的有效性。通过重用静态视觉-文本上下文并更新紧凑的循环状态,SSVR支持高效的多步推理,在预构建前缀缓存的情况下,其Maze解码滚动速度比评估的基线快达98.58倍。

英文摘要

Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to $98.58\times$ faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.

发表机构

  • FDU(复旦大学)
  • HUST(华中科技大学)
  • TJU(天津大学)
  • CCNU(华中师范大学)
  • Kuaishou(快手)

机构由 AI 辅助整理,请以论文原文为准。

↑