arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33269cs.ROcs.CV

Q-WAM:基于动作子空间保护的世界动作模型4比特量化

Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection

Arash Akbari, Arman Akbari, Jingwu Luo, Yuhao Lei, Yi Gao, Weiwei Chen, Xuan Zhang, Zhenman Fang, Geng Yuan, Yanzhi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出Q-WAM,一种针对世界动作模型的4比特量化方法,通过动作可观测性格拉姆矩阵和动作子空间保护保留关键动作信息,在RoboTwin基准上达到接近16比特性能,内存减少3倍以上。

中文摘要 AI 辅助

世界动作模型(WAMs)通过迭代扩散联合生成视频和机器人动作,并在机器人操作中表现出色。然而,其高昂的计算和内存成本带来了巨大的部署挑战。训练后量化(PTQ)可以降低这些成本,但现有的PTQ方法(如平滑和旋转)不足以保持动作生成的精度。为克服这一局限,我们提出Q-WAM,一种针对WAMs的新型4比特权重-激活量化方法,能够保留模型生成的动作。具体而言,我们引入了动作可观测性格拉姆矩阵(AOG),用于衡量层输入通道的每个加权组合中的舍入误差通过所有去噪步骤对最终动作的影响程度。我们还开发了动作子空间保护(ASP),将少数最敏感的动作通道组合保留在一个微小的16比特低秩分支中,并将互补的权重和激活量化为4比特,两者均以密集矩阵乘法形式在GPU上高效运行。最后,为了以最小开销保持动作质量,我们通过聚合每个专家各层的AOG导出的动作质量,识别对生成动作最重要的专家,并仅对这些专家应用ASP。我们在三个WAMs上评估了Q-WAM,包括仿真和实际部署。在RoboTwin 2.0基准上,其平均成功率达到89.6%至93.0%,与16比特模型相差在1.1个百分点以内,同时将量化块的内存减少了3.1至3.4倍。我们的方法在性能上比最强基线SVDQuant高出2.5至8.7个百分点。在宇树G1人形机器人和双臂UR3机器人上,其成功率比SVDQuant提高了12.8至17.6个百分点。

英文摘要

World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but existing PTQ methods such as smoothing and rotation are insufficient to maintain the precision of action generation. To overcome this limitation, we propose Q-WAM, a new 4-bit weight-activation quantization for WAMs that preserves the actions the model generates. Specifically, we introduce the \textit{Action Observability Gramian (AOG)}, which measures how much rounding errors in each weighted combination of a layer's input channels change the final action through all denoising steps. We also develop Action-Subspace Protection (ASP), which keeps the few most action-sensitive channel combinations in a tiny 16-bit low-rank branch and quantizes the complementary weights and activations to 4 bits, both as dense matrix multiplications that run efficiently on GPUs. Finally, to preserve action quality with minimal overhead, we identify the experts that matter most for the generated action by aggregating the AOG-derived action mass across the layers of each expert and apply ASP only to those experts. We evaluate Q-WAM on three WAMs, both in simulation and in real-world deployment. On the RoboTwin 2.0 benchmark, it reaches 89.6--93.0\% average success rate, within 1.1 percentage points of the 16-bit models, while reducing the memory of the quantized blocks by 3.1--3.4$\times$. Our method outperforms the strongest baseline, SVDQuant, by 2.5--8.7 percentage points. On a Unitree G1 humanoid and a bimanual UR3 robot, it improves success over SVDQuant by 12.8-17.6 percentage points.

↑