arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Reflex:通过流推理实现实时VLA控制

Reflex: Real-Time VLA Control through Streaming Inference

Yuanchun Guo, Bingyan Liu

arXiv 2607.14695首次发表:更新:

AI 中文总结

研究针对流匹配VLA模型与实时机器人技术的不兼容问题,提出Reflex框架,利用时间步不变性属性实现实时流推理,通过分区缓存更新、引入自适应归一化层、异步管道和算子融合等方法,在基准测试中实现推理加速和稳定流,减少反应延迟。

AI 中文摘要

流匹配视觉语言动作(VLA)模型有望实现精确的连续控制,但其迭代去噪性质与实时机器人技术存在根本不兼容。我们提出了Reflex框架,通过利用时间步不变性属性,实现了流匹配策略的实时流推理。该框架将注意力上下文分为静态、滑动和动态区域,实现了O(1)增量缓存更新。我们还引入了AdaRMSNorm自适应归一化层,通过异步管道和算子融合进一步提高了吞吐量。在LIBERO和Kinetix基准测试中,Reflex实现了2.58倍的推理加速和50Hz的稳定流,减少了高达54%的反应延迟。

英文摘要

Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow $O(N^2)$ re-computation or mathematically incorrect cache reuse. We present \textbf{Reflex}, a framework that enables \textit{real-time streaming inference} for flow matching policies by exploiting the \textit{Timestep-Invariance Property} -- that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling $O(1)$ incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce \textit{AdaRMSNorm}, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an \textit{async pipeline} that decouples visual encoding from action generation, combined with \textit{operator fusion} that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58$\times$ inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54\% and enabling efficient deployment without performance degradation.

CommentsAccepted by ICML2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑