arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Jetson-PI:通过前瞻对齐异步推理实现机载实时机器人控制

EagleVLA: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Zebin Yang, Qi Wang, Yunhe Wang, Xiurui Guo, Bo Yu, Shaoshan Liu, Jiafeng Xu, Hao Dong, Meng Li

arXiv 2607.12659首次发表:更新:

发表机构

Peking University; AIRS; PrimeBot Research Institute(北京大学; 无合适中文名,保留英文; PrimeBot 研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对低功耗机载设备部署VLA模型的挑战,提出Jetson-PI方法。通过训练轻量级未来校正模块解决感知-执行未对齐,引入基于置信度的调度优化减少反应时间,辅以系统级加速,实验显示其显著提升控制频率和成功率。

AI 中文摘要

视觉-语言-动作(VLA)模型在各种具体任务中取得了令人瞩目的性能。然而,由于其高计算复杂度,在如Jetson Orin这样的低功耗机载设备上部署VLA模型仍具有挑战性,会导致显著的推理延迟和低控制频率。异步推理可通过并行化动作执行和后续推理部分掩盖此延迟,但引入了感知-执行未对齐和长反应时间两个关键问题。本文提出Jetson-PI,一种通过前瞻对齐异步校正实现机载设备上高效VLA部署的方法。为解决未对齐问题,训练了一个轻量级未来校正模块,以预测基于已执行动作的未来环境表示,使动作专家能直接从未来时间步预测动作。为减少反应时间,引入基于置信度的调度优化,自适应平衡VLM和动作专家调用,并辅以系统级加速,包括CUDA图重用、GPU驻留中间缓冲和流展开。大量实验表明,Jetson-PI在NVIDIA Jetson Orin上的控制频率比朴素PyTorch和另一方法分别提高了8.66倍和5.41倍,在LIBERO基准测试中的平均成功率比VLASH高14.8%。

英文摘要

Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational complexity, which leads to substantial inference latency and low control frequency. Asynchronous inference can partially mask this latency by parallelizing action execution and subsequent inference, but it introduces two critical issues: perception-execution misalignment and long reaction time. In this paper, we propose EagleVLA, a method for efficient VLA deployment on onboard devices via Foresight-Aligned Asynchronous Correction. To address misalignment, we train a lightweight future correction module that predicts future environment representations, allowing the action expert to predict actions from the future time step. To reduce reaction time, we introduce confidence-based scheduling optimization that adaptively balances VLM and action expert invocations. We also build a llama.cpp-based inference engine tailored for onboard VLA deployment, with system-level optimizations including CUDA graph reuse, GPU-resident intermediate buffering, and flow unrolling. Extensive experiments demonstrate that EagleVLA achieves 9.85x and 6.19x improvements in control frequency compared with naive PyTorch and vla.cpp on Jetson Orin, while outperforming VLASH by 14.8% in success rate on LIBERO benchmark. The code of our asynchronous algorithm is available on https://github.com/PKU-SEC-Lab/EagleVLA, and our efficient llama.cpp-based inference engine is available on https://github.com/PKU-SEC-Lab/EagleVLA-Edge.

CommentsCoRL 2026 (Spotlight). Formerly known as "Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference"

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑