arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TurboVLA:在RTX 4090上以32 Hz运行、显存占用<1 GB的实时视觉-语言-动作模型

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hengyi Xie, Chenfei Yao, Xianjin Wu, Yingying Zhu, Dingkang Liang, Xiang Bai, Han Ding

arXiv 2607.27205首次发表:更新:

发表机构

Huazhong University of Science and Technology; Huawei Technologies Co. Ltd(华中科技大学; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出TurboVLA,将传统VLA模型的LLM中心路径重构为视觉语言直接映射,仅0.2B参数即可在RTX 4090上实现97.7%LIBERO任务成功率,性能优于更大模型且显存占用极低。

AI 中文摘要

视觉-语言-动作(VLA)模型通常采用以大语言模型(LLM)为中心的$V \to L \to A$路径,即先将视觉观测投影到大语言模型的表示空间,再解码为机器人动作。尽管该设计有效,但每次策略调用都会产生大量计算与内存开销。本研究提出TurboVLA,一种新型VLA范式,将传统$V \to L \to A$路径重新构建为直接的$V + L \to A$映射。TurboVLA不使用大语言模型作为感知与动作间的核心接口,而是独立编码视觉观测与语言指令,通过轻量双向视觉-语言交互直接交换两者信息,并通过紧凑解码器预测连续动作块。这种简洁设计直接从视觉和语言特征构建任务条件表示,大幅降低VLA推理的计算与内存成本。在LIBERO数据集上,TurboVLA仅用0.2B参数,在消费级RTX 4090上实现97.7%的平均成功率、31.2 ms的推理延迟和0.9 GB的推理显存占用,性能匹配或显著优于更大规模的VLA策略。这些结果表明TurboVLA是当前主流以LLM为中心的VLA范式的简洁有效替代方案,为视觉、语言与动作如何关联以实现高效机器人操纵提供了新视角。代码可在此https URL获取。

英文摘要

Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, processing visual observations and language instructions through a large language model before predicting robot actions. Although effective, this design incurs substantial computation and memory overhead. In this work, we introduce TurboVLA, a compact VLA architecture built on a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design directly constructs task-conditioned representations while avoiding the overhead of an LLM-centered execution pathway. On LIBERO, TurboVLA achieves 97.6% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GiB inference VRAM on a consumer-grade RTX 4090. Notably, a 0.4B TurboVLA achieves 88.06% success on RoboTwin 2.0, even matching or outperforming substantially larger VLA policies. These results demonstrate that the simple $V + L \to A$ design of TurboVLA can achieve high performance without requiring an LLM-centric execution pathway, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.

CommentsCode is available at https://github.com/H-EmbodVis/TurboVLA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑