arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QuantWAMs:为世界动作模型校准至合适粒度

QuantWAMs: Calibrating at the Right Granularity for World Action Models

Jiacheng Zhou, Jinfan Lv, Ruixuan Li, Longtai Zhang, Yan Wang, Wenqiang Zhang, Lizhe Qi

arXiv 2607.28405首次发表:更新:

发表机构

College of Intelligent Robotics and Advanced Manufacturing, Fudan University; School of Data Science and Engineering, East China Normal University(复旦大学智能机器人与先进制造学院; 华东师范大学数据科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出QuantWAMs框架,通过三种校准策略优化后训练量化,在WAMs上实现内存大幅降低与速度提升,同时保持与FP16接近的性能,验证了部署可行性。

AI 中文摘要

世界动作模型(World Action Models,WAMs)可联合预测未来观测与动作,但其迭代去噪与闭环执行导致高效部署成本高昂。现有后训练量化(Post-Training Quantization,PTQ)方法不适用于WAMs,因为它们依赖开环目标、同构模型假设及无法反映部署场景的校准分布。本文提出QuantWAMs,一种将量化决策与模型结构、滚动分布及任务目标定义的校准上下文对齐的PTQ框架。QuantWAMs引入三种策略:共享基离群值校准,仅在坐标兼容模块间汇集激活证据;协同训练目标显著性,从联合视频-动作梯度计算经验费舍尔分数,并在校准稳定层粒度分配权重精度;固定干预滚动审计,利用可达闭环状态修正去噪步保护调度,且不改变精度预算。我们在Fast-WAM和LingBot-VA模型上,于RoboTwin 2.0、LIBERO基准及搭载AgiBot G2的真实机器人操纵场景中评估QuantWAMs。在W4A4主导设置下,报告的仿真均值与FP16的差异为0.2至0.7个百分点。真实机器人试验进一步验证了其在三项操纵任务上的部署可行性。针对目标视频和动作块,QuantWAMs将峰值权重与激活内存降至FP16的约29%,并提供1.4至1.6倍的块级加速。

英文摘要

World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-training quantization (PTQ) methods are poorly suited to WAMs because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment. We present QuantWAMs, a PTQ framework that aligns quantization decisions with the calibration context defined by model structure, rollout distribution, and task objective. QuantWAMs introduces three strategies: shared-basis outlier calibration, which pools activation evidence only across coordinate-compatible modules; co-training-objective saliency, which computes empirical-Fisher scores from the joint video--action gradient and assigns weight precision at a calibration-stable layer granularity; and fixed-intervention rollout auditing, which revises denoising-step protection schedules using reachable closed-loop states without changing the precision budget. We evaluate QuantWAMs on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and real-robot manipulation with an AgiBot G2. Under a W4A4-dominant setting, the reported simulation means differ from FP16 by 0.2--0.7 percentage points. Real-robot trials further establish deployment feasibility on three manipulation tasks. For the targeted video and action blocks, QuantWAMs reduces peak weight-and-activation memory to about 29\% of FP16 and provides 1.4--1.6$\times$ block-level speedups.

Comments13 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑