解决状态表示不匹配:面向开环VLA规划的状态空间视觉推理
Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning
浏览论文内容
中文总结 AI 辅助
针对开环VLA规划中的状态表示不匹配问题,提出SSVR方法,通过解耦静态上下文与循环潜在状态,实现高效多步推理,显著提升规划性能。
中文摘要 AI 辅助
尽管视觉-语言-动作(VLA)模型取得了快速进展,现有推理范式在开环规划中仍面临根本性的“状态表示不匹配”问题。在仅给定初始观测的情况下,模型必须在内部模拟动作条件下的状态转移,而文本、像素和潜在空间推理分别可能遭受有损空间压缩、误差累积的视觉生成以及中间潜在标记的绕过,从而削弱了可靠的长时程规划能力。我们提出状态空间视觉推理(SSVR),该方法将静态视觉上下文、语言约束和循环潜在状态解耦。SSVR对初始图像和指令仅编码一次,然后基于潜在状态调节每个动作预测,并通过动作条件下的GRU更新该状态。以Qwen2.5-VL作为骨干模型,SSVR在FrozenLake、Maze和MiniBehavior上分别达到99.5/99.6、96.3/98.0和83.9/90.6的EM/PR,显著优于先前方法。大量实验支持了循环状态建模在VLA开环规划中跨输入变换和迁移设置的有效性。通过重用静态视觉-文本上下文并更新紧凑的循环状态,SSVR支持高效的多步推理,在预构建前缀缓存的情况下,其Maze解码滚动速度比评估的基线快达98.58倍。
英文摘要
Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to $98.58\times$ faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.
发表机构
- FDU(复旦大学)
- HUST(华中科技大学)
- TJU(天津大学)
- CCNU(华中师范大学)
- Kuaishou(快手)
机构由 AI 辅助整理,请以论文原文为准。