arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于高效视觉-语言-动作推理的具有质心重用的运动感知向量量化框架

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

Zhuoran Song, Haozhe Jiang, Chunyu Qi, Minnan Pei, Gang Li, Xiaoyao Liang, Haibing Guan

arXiv 2607.24148首次发表:更新:

发表机构

School of Computer Science, Shanghai Jiao Tong University; Institute of Automation, Chinese Academy of Sciences(上海交通大学计算机科学与工程系; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型推理延迟高问题,提出VQVLA算法-硬件协同设计框架,利用MotionVQ动态调整量化精度,采用合并质心向量量化通用矩阵乘法范式,设计加速器,实验表明其相比多种方法有显著加速且精度损失小。

AI 中文摘要

视觉-语言-动作(VLA)模型在具身人工智能中展现出强大潜力,但其在GPU上的推理延迟高限制了实时部署。现有加速器虽提高了效率,但将VLA模型视为全精度工作负载,未充分利用内存和计算中的大量冗余。本文提出VQVLA,一种算法-硬件协同设计框架,通过利用权重相似性和执行动态性加速VLA推理。先引入MotionVQ,一种基于机器人执行状态动态调整量化精度的运动感知向量量化方案,在保持任务成功率的同时减少内存访问。然后提出合并质心向量量化通用矩阵乘法范式,通过质心的空间聚合和时间重用消除冗余乘法。为实现这些优化,设计了一个有效支持动态精度选择和质心重用计算的加速器。实验结果表明,VQVLA分别比A100 GPU、Dadu-Corki、LUT-DLA、CodeGEMM和ShiftAddLLM加速6.5倍、2.8倍、1.9倍、3.3倍和4.3倍,精度下降可忽略不计。

英文摘要

Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑