发表机构
VinRobotics; Center for AI Research, VinUniversity; Intelligent Autonomous Systems, TU Darmstadt(VinRobotics; VinUniversity人工智能研究中心; 达姆施塔特工业大学智能自主系统)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出vla.simd CPU推理引擎和IMPACT策略,实现语言条件操控的高效CPU部署,在多个CPU上加速1.4倍,并在Raspberry Pi 5上达到33.5动作/秒,GPU评估成功率76.4%。
AI 中文摘要
在没有专用GPU的情况下部署语言条件操控,需要高效的推理以及能够覆盖策略查询之间延迟的动作块。我们提出了vla.simd,一个CPU推理引擎,它结合了共享的SIMD微内核、可复用计算和目标特定优化。我们将查询延迟和执行范围与滞后和时间对齐执行下的动作可用性联系起来,区分动作供应与反馈频率。在六个策略和四个CPU上,vla.simd相对于编译的PyTorch参考实现实现了约1.4倍的中位加速,同时保持了fp32数值保真度。我们还引入了IMPACT,一种基于ACT的策略,具有缓存的文本表示和语言调制的视觉特征。IMPACT是我们评估集中唯一在Raspberry Pi 5上提供至少30个动作/秒的语言条件策略:经过90秒热浸泡后,它在fp32下提供33.5个动作/秒,在int8下提供81.2个动作/秒。单独的GPU评估在四个LIBERO套件上实现了76.4%的平均成功率,无需机器人预训练;指令打乱测试证明了对熟悉目标的选择能力。在SO-101机械臂上使用IMPACT和在UR10e上使用带Robotiq夹爪的SmolVLA的试验,展示了在两种机器人实体上的CPU部署。
英文摘要
Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately $1.4\times$ median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield $76.4\%$ mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.
Comments8 pages, 7 tables, 5 figures. Project page: https://vla-simd.github.io/