arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39056cs.RO

SteerQuant:在世界-动作模型中通过动作引导缩放来引导量化误差

SteerQuant: Steering Quantization Error with Action-Guided Scaling in World-Action Models

  • Tongji(同济大学)
  • HKUST(香港科技大学)
  • HIT(哈尔滨工业大学)
  • EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yunhan Wang, Haodong Wang, Zhiming Liu, Zicong Hong, Qianli Liu, Xiaoyi Pang, Yangjia Hu, Quanxin Shou, Yikun Miao, Song Guo

中文总结 AI 辅助

针对世界-动作模型,提出SteerQuant 4比特量化框架,通过动作引导缩放将量化误差导向低影响计算,并开发Rudder推理引擎,在保持成功率的同时实现显著加速。

中文摘要 AI 辅助

世界-动作模型(WAMs)通过迭代去噪联合生成未来的世界状态和动作,使用共享权重处理视频、本体感觉和动作标记的异构语义流。量化降低了推理成本,但不同流中相当的数字误差可能对最终动作产生显著不同的影响,使得仅凭数值精度不足以实现可靠控制。我们提出了SteerQuant,一个针对WAMs的4比特量化框架,它将误差引导到对最终动作影响较小的计算上。它映射了每个流的量化误差如何影响最终动作,并利用该映射指导共享通道缩放。激活缩放进一步针对每个流和去噪步骤进行校准,以适应激活范围和动作影响的变化。这使得量化适应不同流的需求,而无需复制权重或增加选定流的比特宽度。为了减少缩放引入的额外内核启动和内存流量,我们开发了Rudder,一个针对WAMs的4比特推理引擎,它将缩放和输出补偿融合到低比特内核中。在W4A8和W4A4下,SteerQuant将平均LIBERO成功率保持在全精度的0.8个百分点以内,同时在三个WAMs上相比BF16提供了高达2.23倍的去噪加速,并降低了峰值GPU内存使用。在真实的双臂机器人上,W4A8部署实现了1.35倍的端到端推理加速,同时相对于BF16保持了平均任务成功率。

英文摘要

World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly different effects on final actions, making numerical accuracy alone insufficient for reliable control. We introduce SteerQuant, a 4-bit quantization framework for WAMs that steers errors toward computations with less influence on final actions. It maps how each stream's quantization errors affect final actions and uses this map to guide shared channel scaling. Activation scaling is further calibrated for each stream and denoising step to accommodate changes in activation ranges and action impact. This adapts quantization to different stream requirements without duplicating weights or increasing bit-widths for selected streams. To reduce the extra kernel launches and memory traffic introduced by scaling, we develop Rudder, a 4-bit inference engine for WAMs that fuses scaling and output compensation into low-bit kernels. Under W4A8 and W4A4, SteerQuant maintains mean LIBERO success within 0.8 percentage points of full precision, while delivering up to $2.23\times$ denoising speedup over BF16 across three WAMs with reduced peak GPU memory usage. On a real dual-arm robot, W4A8 deployment achieves a $1.35\times$ end-to-end inference speedup while maintaining average task success relative to BF16.

补充信息

↑