arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24433cs.ROcs.AIcs.SYeess.SY

FoldQuantVLA:通过一致性折叠实现视觉-语言-动作模型的原生低比特量化

FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen, Thanh Q. Duong, Ngan Le, Meng Guo, Vien A. Ngo, An T. Le

首次发表
浏览论文内容

中文总结 AI 辅助

FoldQuantVLA通过一致性折叠实现视觉-语言-动作模型的训练后低比特量化,在保持机器人行为的同时,以W4A4配置在Orin和桌面GPU上获得1.2-1.52倍加速,并将真实任务成功率从80%提升至92.5%。

中文摘要 AI 辅助

低比特视觉-语言-动作推理必须在保持机器人行为的同时降低从观察到动作的延迟。我们提出了FoldQuantVLA,一个训练后量化框架,在标定、权重舍入和原生整数执行过程中保持一致的激活表示。它结合了通道缩放和块Hadamard变换与动态逐token量化,无需策略重训练。自定义TensorRT插件在语言主干和迭代动作专家中执行投影,在Ada GPU和Jetson AGX Orin上使用四比特权重和激活(W4A4)。评估涵盖LIBERO、SimplerEnv和两个机器人平台。在三个GR00T检查点和$\pi_{0.5}$上,W4A4在Orin上相比浮点TensorRT实现了$1.20$到$1.33\times$的加速,在桌面上实现了$1.25$到$1.52\times$的加速。将语言注意力输出和前馈下投影保留为八比特(W8A8)在所有四个检查点上提高了保留动作保真度。在四个真实机器人任务中,该配置将观察到的GR00T N1.7成功率从均匀W4A4的$80.0\\%$提高到每个配置80次试验中的$92.5\\%$,并测得Orin额外延迟为1毫秒。

英文摘要

Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and $π_{0.5}$, W4A4 achieves $1.20$ to $1.33\times$ speedups over floating-point TensorRT on Orin and $1.25$ to $1.52\times$ on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from $80.0\%$ with uniform W4A4 to $92.5\%$ over 80 trials per configuration, with a measured additional Orin latency of 1 ms.

发表机构

  • VinRobotics
  • University of Arkansas(阿肯色大学)
  • Peking University(北京大学)
  • VinUniversity
  • TU Darmstadt(达姆施塔特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑