arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Deltoris:通过比特级稀疏性和推测性推理实现具身AI中的实时VLA推理

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

Zheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He, Yinhe Han, Jingwen Leng, Minyi Guo, Yiming Gan, Yu Feng

arXiv 2608.04428首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Qi Zhi Institute; ICT, Chinese Academy of Sciences(上海交通大学; 上海期智研究院; 中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Deltoris作为算法-硬件协同设计框架,通过时间感知比特级稀疏性、推测性推理及定制加速器,实现具身AI中实时VLA推理,加速比显著且精度相当。

AI 中文摘要

视觉-语言-动作(VLA)模型已成为具身AI的关键组件,现有方法中,基于扩散的VLA模型具备优异的运动质量与泛化能力,但这类模型计算密集,需以50-200Hz的高控制频率运行,对边缘设备施加了严格的延迟与功耗约束。本研究提出Deltoris——一款面向高效扩散型VLA推理的算法-硬件协同设计框架:首先,利用连续输入的时间相似性,提出时间感知比特级稀疏性算法,仅计算连续输入间的差异,消除冗余的比特级操作;其次,为解决该算法引入的额外片外数据传输,提出推测性推理技术,将数据加载分摊至多步控制过程;最后,为支撑上述技术,协同设计了一款专用加速器,配备定制的1D脉动比特级串行处理单元(PE)阵列,消除处理单元负载不均衡问题。评估显示,Deltoris相较移动GPU实现了最高34.2倍的加速,相较现有加速器实现了6.1倍加速,同时保持了相当的精度。

英文摘要

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑