发表机构
Shanghai Jiao Tong University; Shanghai Qi Zhi Institute; ICT, Chinese Academy of Sciences(上海交通大学; 上海期智研究院; 中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Deltoris作为算法-硬件协同设计框架,通过时间感知比特级稀疏性、推测性推理及定制加速器,实现具身AI中实时VLA推理,加速比显著且精度相当。
AI 中文摘要
视觉-语言-动作(VLA)模型已成为具身AI的关键组件,现有方法中,基于扩散的VLA模型具备优异的运动质量与泛化能力,但这类模型计算密集,需以50-200Hz的高控制频率运行,对边缘设备施加了严格的延迟与功耗约束。本研究提出Deltoris——一款面向高效扩散型VLA推理的算法-硬件协同设计框架:首先,利用连续输入的时间相似性,提出时间感知比特级稀疏性算法,仅计算连续输入间的差异,消除冗余的比特级操作;其次,为解决该算法引入的额外片外数据传输,提出推测性推理技术,将数据加载分摊至多步控制过程;最后,为支撑上述技术,协同设计了一款专用加速器,配备定制的1D脉动比特级串行处理单元(PE)阵列,消除处理单元负载不均衡问题。评估显示,Deltoris相较移动GPU实现了最高34.2倍的加速,相较现有加速器实现了6.1倍加速,同时保持了相当的精度。
英文摘要
Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.