发表机构
The University of Tokyo; Institute of Science Tokyo(东京大学; 东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VLA-ULAP框架,将云端VLA调用与超轻量级本地动作预测器交错执行,在边缘设备上大幅降低延迟和能耗,同时保持高成功率。
AI 中文摘要
十亿参数级别的视觉-语言-动作(VLA)策略需要大量的机载功耗,而远程推理中的通信延迟阻碍了及时响应。我们提出了VLA-ULAP,它将远程VLA调用与超轻量级本地动作预测器(ULAP)交错执行。ULAP包含约740万参数(包括冻结的视觉编码器),结合当前视图、本体感觉和执行动作历史,一次性预测动作块。ULAP独立训练,不需要VLA隐藏状态、在线验证或服务器往返。在Jetson Orin Nano上,ULAP每次推理耗时19.9毫秒、能耗0.183焦耳,而GR00T在RTX A6000上耗时284.3毫秒、能耗50.55焦耳。在三个模拟基础策略/基准测试组合中,选定的工作点移除了48.8%至76.7%的VLA调用,同时保留了95.0%至97.5%的基线成功率。与VLA-JEPA上的本地VLA加速替代方案相比,在相当成功率下,ULAP相比ACT每个成功回合估计减少49.2%的推理时间和51.0%的GPU能耗;在相同成功率下,相比SP-VLA减少77.1%的时间和79.9%的能耗。物理SO-101实验在已见和未见放置场景中保留了95.2%至100%的基线成功率,同时基于成功回合调用次数和实测设备成本,估计推理时间减少47.9%至58.0%,推理设备能耗减少52.1%至62.5%。更快的响应还提高了动态任务成功率:在考虑延迟的LIBERO-Safety仿真中,VLA-ULAP在两个任务上超过π0.5达11.0和15.5个百分点,同时VLA调用大约减半。
英文摘要
Billion-parameter vision-language-action (VLA) policies run either onboard, consuming substantial power, or on remote servers, adding communication latency. To address these drawbacks and better balance latency and onboard energy consumption, we propose VLA-ULAP. It partitions inference across decision times, interleaving remote VLA calls with predictions from an Ultra-Lightweight Local Action Predictor (ULAP). A single ULAP has $\sim$7.4M parameters including the frozen vision encoder. It combines current views, proprioception, and executed action history to predict chunks in one pass. Trained independently, it requires no VLA hidden states, online verification, or server round trips. On Jetson Orin Nano, ULAP takes 20.7 ms and 0.122 J of idle-subtracted energy per inference, versus 289.3 ms and 40.46 J for GR00T on RTX A6000. Across four simulated base-policy/benchmark pairs, VLA-ULAP removes 45.3-77.3% of VLA calls while retaining 95.0-98.5% of baseline success rates at selected operating points. On VLA-JEPA, VLA-ULAP also surpasses local acceleration alternatives, using an estimated 49.4% less inference time and 51.5% less GPU energy per successful episode than ACT, and 77.5% less time and 80.7% less energy than SP-VLA at higher success rates. In physical SO-101 trials, it similarly removes 70.3-71.9% of VLA calls without observed success-rate loss at seen or held-out placements. Measured device costs imply 64.6-66.4% less inference time and 70.1-71.7% less idle-subtracted energy per successful episode at these call counts. Beyond these savings, faster responses help VLA-ULAP exceed $π_{0.5}$'s success rate by 11.0 and 15.5 percentage points (pp) on two tasks in latency-aware LIBERO-Safety simulation while approximately halving VLA calls.
CommentsPreprint