arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CHASE-VLA:面向视觉-语言-动作模型的块感知尺度估计后训练量化框架

CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation

Jin Hyun, Jung Gyu Min, Gyuhyun Jung, Youngjoo Lee

arXiv 2610.02666首次发表:更新:

发表机构

KAIST; POSTECH(韩国科学技术院; 浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CHASE-VLA,利用动作块上下文与去噪步组信息自适应激活尺度,实现VLA模型AE的W4A4量化,在LIBERO上恢复FP16性能并大幅降低存储与内存流量。

AI 中文摘要

视觉-语言-动作(VLA)模型将视觉观察和语言指令映射为连续的机器人动作,但基于扩散的动作专家(AE)对低位后训练量化(PTQ)构成了关键挑战。AE在去噪步骤和策略查询中被反复调用,固定的校准尺度可能与随去噪进度和预期运动而变化的激活范围不匹配。我们提出CHASE-VLA,一种块感知的PTQ方法,利用策略中现成的VLA特定信号:生成的动作块,包括其未执行的未来后缀。CHASE-VLA不仅依赖AE层的静态尺度匹配,还将先前生成的块作为因果动作上下文与去噪步骤组信息结合,以调整AE激活尺度。这使得在重复的AE中,MLP和注意力投影均可实现W4A4量化,而无需修改预训练策略。在LIBERO上,当AE中的MLP和注意力投影均量化为W4A4时,CHASE-VLA在π0.5上达到97.3%的平均成功率,恢复了FP16级性能。CHASE-VLA还将量化AE线性层的权重存储减少了73.4%,并在π0.5和GR00T N1.6上分别将单块内存流量减少了70.9%和71.2%,预测器开销最多为节省存储的1.26%。

英文摘要

Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration scales can be mismatched with activation ranges that vary with denoising progress and intended motion. We propose CHASE-VLA, a chunk-aware PTQ method that exploits a VLA-specific signal readily available from the policy: the generated action chunk, including its unexecuted future suffix. Rather than relying only on static scale matching for AE layers, CHASE-VLA combines the previously generated chunk as causal action context with denoising step group information to adapt AE activation scales. This enables W4A4 quantization of both MLP and attention projections in the repeated AE without modifying the pretrained policy. On LIBERO, CHASE-VLA achieves 97.3% average success rate on $π_{0.5}$ when both MLP and attention projections in the AE are quantized to W4A4, restoring FP16-level performance. CHASE-VLA also reduces the weight storage of the quantized AE linear layers by 73.4% and their single-chunk memory traffic by 70.9% and 71.2% on $π_{0.5}$ and GR00T N1.6, respectively, with a predictor overhead of at most 1.26% of the saved storage.

CommentsAccepted at ACCV 2026. 22 pages, including references and supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑