发表机构
Stevens Institute of Technology; Northeastern University; Arizona State University; University of Georgia(史蒂文斯理工学院; 东北大学; 亚利桑那州立大学; 佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出IG-VLA框架,通过潜在时空推理和场景要点记忆内化未来想象,在不增加推理开销的情况下提升VLA机器人操作性能,在LIBERO-Plus上成功率提升近6%,推理速度最高提升6.38倍。
AI 中文摘要
视觉-语言-动作(VLA)模型越来越多地引入中间推理以改进机器人操作,然而现有方法主要对观察到的状态进行推理,并未显式预测未来的场景演变。然而,在每一步推理中都将这种推理扩展到显式的未来推演,会带来大量的计算开销。我们提出IG-VLA,一个VLA推理框架,使模型能够想象未来并内化要点。我们的潜在时空推理学习直接在视觉表示空间中想象与任务相关的未来场景演变,从而在没有昂贵的像素级视频生成的情况下指导动作预测。为了进一步减少推理开销,我们引入了场景要点记忆,它将推理得出的场景-行为关联内化到一个紧凑的场景要点标记中,保留了未来推理的好处,同时在推理时绕过了显式的未来想象。在LIBERO、LIBERO-Plus和VLABench上的大量实验证明了IG-VLA的有效性和效率。在LIBERO-Plus语言套件上,推理策略和要点策略在成功率上均比最强基线高出近6%。要点策略还实现了比基线最高6.38倍的加速,在单个NVIDIA A6000 GPU上将每个动作块的推理延迟从1081毫秒降低到169.5毫秒。这些结果表明,未来的时空推理可以被有效地内化,以实现高效的VLA部署。
英文摘要
Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, however, introduces substantial computational overhead. We propose IG-VLA, a VLA reasoning framework that enables models to imagine the future and internalize the gist. Our Latent Spatiotemporal Reasoning learns to imagine task-relevant future scene evolution directly in visual representation space, guiding action prediction without costly pixel-level video generation. To further reduce inference overhead, we introduce Scene Gist Memory, which internalizes reasoning-derived scene-behavior associations into a compact Scene Gist Token, preserving the benefits of future reasoning while bypassing explicit future imagination at inference. Extensive experiments on LIBERO, LIBERO-Plus, and VLABench demonstrate the effectiveness and efficiency of IG-VLA. On the LIBERO-Plus Language suite, both the reasoning and gist policies outperform the strongest baseline by nearly 6% in success rate. The gist policy also achieves up to 6.38x speedup over baselines, reducing inference latency from 1081ms to 169.5ms per action chunk on a single NVIDIA A6000 GPU. These results demonstrate that future spatiotemporal reasoning can be effectively internalized for efficient VLA deployment.