arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03959cs.AI

LatentQuant:在NVFP4 VAE量化下保持面向策略的潜在契约

LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization

  • UC San Diego(加州大学圣迭戈分校)
  • Sun Yat-Sen University(中山大学)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • HKUST (Guangzhou)(香港科技大学(广州))
  • XSquare Robot(XSquare机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

Ziye Deng, Lufang Chen, Shuyu Feng, Zhenwei Duan, Zicong Ye, Yu Sun, Xiaofan Li, Ruyi Gan, Hao Wang, Hao Zhang

AI总结:

针对NVFP4量化破坏策略潜在契约的问题,提出两阶段QAT框架LatentQuant,先对齐编码器再冻结并适配解码器,在保持重建质量的同时恢复控制性能。

AI中文摘要:

近期世界动作模型(WAMs)复用预训练的视频VAE,其编码器潜在表示直接条件化下游动作策略。因此,量化不仅需要保持重建保真度,还必须保持冻结策略所期望的面向策略的潜在契约。直接NVFP4量化未补偿W4A4量化误差,而联合量化感知训练(QAT)可以通过移动该表示来恢复重建。在Wan2.1上,联合QAT几乎匹配FP32 VBench-7(0.7403对比0.7409),但LIBERO成功率从95.5%骤降至10.5%。受控的仅解码器实验表明,激活量化-反量化操作改变了重建信号和解码器雅可比矩阵,将返回给编码器的梯度重定向,并诱发持续的潜在漂移。基于该机制,我们提出LatentQuant,一种两阶段NVFP4 QAT框架,首先将量化编码器与其高精度对应物对齐,然后冻结编码器同时适配解码器。在Wan2.1和Wan2.2上,LatentQuant保持了接近基线的控制和高质量重建,在LIBERO上达到95.75%的成功率,在RoboTwin上达到68.8%。在NVIDIA B300 GPU上,NVFP4执行相比BF16 cuDNN实现了1.17倍至1.26倍的端到端VAE加速。

英文摘要:

Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quantization-aware training (QAT) can recover reconstruction by moving this representation. On Wan2.1, joint QAT nearly matches FP32 VBench-7 (0.7403 versus 0.7409), yet LIBERO success collapses from 95.5% to 10.5%. Controlled decoder-only experiments show that activation quantize-dequantize operations alter the reconstruction signal and decoder Jacobian, redirecting the gradient returned to the encoder and inducing persistent latent drift. Based on this mechanism, we introduce LatentQuant, a two-stage NVFP4 QAT framework that first aligns the quantized encoder with its high-precision counterpart, then freezes it while adapting the decoder. Across Wan2.1 and Wan2.2, LatentQuant preserves near-baseline control and high reconstruction quality, achieving 95.75% success on LIBERO and 68.8% on RoboTwin. On NVIDIA B300 GPUs, NVFP4 execution achieves 1.17x-1.26x end-to-end VAE speedups over BF16 cuDNN.

↑