发表机构
Colorado School of Mines; University of Central Florida; USC Institute for Creative Technologies; Microsoft AI(科罗拉多矿业大学; 中佛罗里达大学; 南加州大学创意技术研究所; 微软人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VLAQuantBench通过409次运行和94,574个仿真回合,系统评估了VLA模型训练后量化,发现层范围、数值格式与校准的交互决定成败,并提出了具体的精度分配方案。
AI 中文摘要
训练后量化降低了视觉-语言-动作(VLA)模型的内存需求,但精度选择必须考虑层范围、数值格式和校准之间的相互作用。我们引入了VLAQuantBench,一个受控评估,包含409次运行和94,574个仿真回合:在LIBERO上评估四个模型,X-VLA额外在三个仿真基准族上评估。在未校准的W4A4四舍五入到最近量化下,将π0.5动作头子集从126层扩展到167层,成功率从7.0%提高到70.5%。固定观测重放确认了相应的数值恢复。两回合校准消除了测试子集中的严重联合故障,而相同的平滑和裁剪方案降低了π0的成功率,且未能恢复OpenVLA-OFT的端到端性能。对于OpenVLA-OFT,保护一个28,672参数的输出投影反而恢复了接近基线的成功率:剩余的441个合格线性层在LIBERO-Long上保留W3,或在所有四个套件中保留八位激活。任务聚类区间支持大的失败和恢复对比。这些结果确立了依赖配方的交互,并确定了具体的精度分配,而非通用的层敏感性规则。真实内核和物理机器人测量补充了准确性分析。代码、配置和回合记录在此https URL公开可用。
英文摘要
Post-training quantization reduces the memory requirements of vision-language-action (VLA) models, but precision selection must account for the interaction between layer scope, numerical format, and calibration. We introduce \textbf{VLAQuantBench}, a controlled evaluation with 409 runs and 94,574 simulation episodes: four models on LIBERO, with X-VLA additionally evaluated on three simulation benchmark families. Under uncalibrated W4A4 round-to-nearest quantization, expanding a $π_{0.5}$ action-head subset from 126 to 167 layers raises success from 7.0\% to 70.5\%. Fixed-observation replay confirms a corresponding numerical recovery. Two-episode calibration removes the severe joint failures in the tested subsets, whereas the same smoothing-and-clipping recipe lowers $π_0$ success and does not recover OpenVLA-OFT end-to-end. For OpenVLA-OFT, protecting one 28,672-parameter output projection instead restores near-baseline success: the remaining 441 eligible linear layers retain W3 on LIBERO-Long or eight-bit activations across all four suites. Task-clustered intervals support the large failure and recovery contrasts. These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules. Real-kernel and physical-robot measurements complement the accuracy analysis. Code, configurations, and episode records are publicly available at https://github.com/jiuyixu25/VLAQuantBench.
Comments28 pages, 35 tables, 4 figures