发表机构
UC San Diego; MIT(加州大学圣迭戈分校; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlashVLA是一种流式动作解码框架,通过分块因果注意力维护动作缓冲,实现低延迟、平稳异步的VLA推理,在单GPU上达≥30Hz控制频率且任务性能优异。
AI 中文摘要
视觉-语言-动作(Vision-Language-Action,VLA)模型在机器人操控领域的应用前景日益广阔,但其实际部署仍受限于高推理延迟与不稳定的异步执行问题。这一挑战在基于流匹配(flow-matching)的VLA模型中尤为突出,此类模型的动作解码需基于视觉语言模型(Vision-Language-Model,VLM)上下文执行多轮迭代步骤。尽管高效推理方法可提升控制频率,异步方法能减少执行空闲时间,但现有方案往往无法同时实现低延迟推理与准确、时间一致的异步执行。本文提出FlashVLA,一种以统一架构解决上述两大挑战的流式动作解码框架。FlashVLA维护包含多个不同噪声水平块的流式动作缓冲,采用分块因果注意力(chunk-wise causal attention)对其解码,该设计使FlashVLA能在每一步推理中生成一个可执行的动作块。此外,其分块自回归架构隐含保留动作连续性,无需额外的未来状态条件即可实现平稳的异步执行。经大量仿真与真实世界实验验证,FlashVLA在维持出色任务性能的同时大幅提升推理速度,在单GPU上可实现≥30Hz的控制频率,且在真实部署中具备平稳的异步推理能力。
英文摘要
Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action decoding requires multiple iterative steps conditioned on the VLM context. While efficient inference methods improve control frequency and asynchronous methods reduce execution idle time, existing approaches often fail to jointly achieve low-latency inference and accurate, temporally consistent asynchronous execution. We introduce \textbf{FlashVLA}, a streaming action decoding framework that addresses both challenges in a unified formulation. FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and decodes them using chunk-wise causal attention. This design allows FlashVLA to produce one executable action chunk per inference step. Moreover, its chunk-wise autoregressive formulation implicitly preserves action continuity, enabling smooth asynchronous execution without extra future-state conditioning. Across extensive simulated and real-world experiments, FlashVLA substantially improves inference speed while maintaining strong task performance. It can achieve $\geq$30\,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.
Comments17 pages, 8 figures