发表机构
Soochow University; Shanghai Jiao Tong University; Yootta; RMIT University(苏州大学; 上海交通大学; Yootta; 皇家墨尔本理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对量化VLA模型泛化能力下降问题,提出脆弱性导向微调框架PIVOT-Q,通过在线策略蒸馏选择性纠正脆弱状态,以7.4%的蒸馏预算实现全面性能恢复。
AI 中文摘要
后训练量化已被证明在标准评估条件下能保持VLA(视觉-语言-动作)模型的性能,但其是否能保持全精度模型的鲁棒性和泛化能力仍未得到充分探索。在本研究中,我们系统地研究了后量化VLA策略在环境扰动下的鲁棒性和泛化能力。实证结果表明,尽管量化策略在分布内性能上保持相当,但它们可能对细微的环境变化变得脆弱。我们进一步观察到,动作差异集中在少量 rollout 状态中,而教师指导根据差异程度产生相反效果:在高差异状态下能提升泛化,但在低差异状态下可能降低泛化。这些发现表明,有效的后量化恢复需要对脆弱状态进行选择性干预,而非全局蒸馏学生模型。因此,我们提出了策略诱导的脆弱性导向微调(PIVOT-Q),一种脆弱性感知的在线策略蒸馏(OPD)框架,该框架利用冻结的全精度策略作为教师,选择性地纠正量化学生 rollout 中遇到的脆弱状态。PIVOT-Q 通过短视界内的折扣累积差异识别脆弱状态,应用相位平衡的稀疏监督,并使用行为锚点防止不必要的更改。在七个 LIBERO-Plus 环境变化下的实验表明,该方法在多个 VLA 骨干网络和量化方法上均能实现一致的恢复。值得注意的是,PIVOT-Q 在所有设置中始终优于全状态蒸馏,同时仅使用其状态级蒸馏预算的 7.4%。我们的代码可在该 https URL 获取。
英文摘要
Post-training quantization has been shown to preserve VLA performance under standard evaluation conditions, but whether it preserves the full-precision model's robustness and generalization remains underexplored. In this study, we systematically study the robustness and generalization of post-quantized VLA policies under environmental disturbances. Empirical results show that quantized policies can become fragile to subtle environmental variations despite retaining comparable in-distribution performance. We further observe that action discrepancies are concentrated in a small subset of rollout states, while teacher guidance has opposite effects depending on discrepancy: it improves generalization at high-discrepancy states but can degrade it at low-discrepancy states. These findings reveal that effective post-quantization recovery requires selectively intervening on vulnerable states rather than globally distilling the student. We therefore propose Policy-Induced Vulnerability-Oriented Tuning (PIVOT-Q), a vulnerability-aware On-Policy Distillation (OPD) framework that selectively corrects vulnerable states encountered during quantized-student rollouts using the frozen full-precision policy as a teacher. PIVOT-Q identifies vulnerable states using discounted accumulated discrepancies over a short horizon, applies phase-balanced sparse supervision, and uses a Behavioral Anchor to prevent unnecessary changes. Experiments under seven LIBERO-Plus environmental variations demonstrate consistent recovery across multiple VLA backbones and quantization methods. Notably, PIVOT-Q consistently outperforms full-state distillation across all settings while using only 7.4% of its state-level distillation budget. Our code is available at https://github.com/ruanruan-andy/PIVOT-Q.