arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAVEL:面向基于流的视觉-语言-动作模型的异步滚动推理

RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

Yuhan Chen, Ke Yu, Pengfei Liu, Shuxun Wang, Yi Yang, Linchao Zhu

arXiv 2609.34170首次发表:更新:

发表机构

Zhejiang University; Nanyang Technological University(浙江大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于流的VLA模型推理延迟高的问题,提出RAVEL异步推理框架,通过滚动缓冲单步去噪和VLM编码解耦,在保持任务能力的同时显著降低响应延迟,实现高频闭环控制。

AI 中文摘要

基于流的视觉-语言-动作(VLA)模型在通用机器人操作中表现出色,但其对计算成本高昂的VLM编码和多步迭代动作生成的依赖,造成了显著的延迟瓶颈。由此产生的推理延迟使得机器人难以快速响应,尤其是在动态环境中。我们通过RAVEL(滚动异步VLA实现低延迟控制)解决了这一局限,这是一种异步推理框架,同时解决了VLM骨干网络和动作专家模块的计算瓶颈。为减少多步动作去噪带来的延迟,RAVEL允许在单步去噪后执行近期动作,方法是将部分去噪的未来动作在滚动缓冲区中向前传递。为避免阻塞于缓慢的VLM编码,RAVEL将VLM编码与滚动动作生成解耦,使动作专家能够利用最新的可用VLM上下文持续运行,同时轻量级快速观测通路(FOP)直接基于当前观测对动作专家进行条件化。在模拟和真实世界的操作任务中,RAVEL持续实现了显著更低的响应延迟,同时保持了底层VLA的任务能力,从而支持高频且响应灵敏的闭环控制。

英文摘要

Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult for robots to respond quickly, especially in dynamic environments. We address this limitation with RAVEL (Rolling Asynchronous VLA Enabling Low-Latency Control), an asynchronous inference framework that addresses the computational bottlenecks of both the VLM backbone and the action expert. To reduce the delay from multi-step action denoising, RAVEL allows near-term actions to be executed after a single denoising step by carrying partially denoised future actions forward in a rolling buffer. To avoid blocking on slow VLM encoding, RAVEL decouples VLM encoding from rolling action generation, allowing the action expert to operate continuously using the latest available VLM context, while a lightweight Fast Observation Pathway (FOP) directly conditions the action expert on current observations. Across simulated and real-world manipulation tasks, RAVEL consistently achieves substantially lower response latency while maintaining the task capability of the underlying VLA, enabling high-frequency and responsive closed-loop control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑