发表机构
Colorado School of Mines; University of Florida; University of Southern California - Institute for Creative Technology(科罗拉多矿业大学; 佛罗里达大学; 南加州大学创意技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PreDE框架,利用离线动作偏差预测量化配置对世界动作模型任务性能的影响,通过校准阈值实现高效配置筛选,实验验证其高覆盖率和准确性。
AI 中文摘要
世界动作模型(WAMs)依赖视频生成骨干网络,部署时需要大量内存和计算资源。训练后量化可减少内存占用并加速推理,但位宽、分组和量化器选择构成了庞大的配置空间。通过穷举闭环评估来识别保持任务性能的配置成本高昂。我们提出PreDE(部署前预测),一个策略校准的框架,用于从离线动作偏差预测量化导致的任务退化。利用小型开发集的闭环结果,PreDE校准两个阈值,并使用固定观测日志接受、拒绝或延迟新配置。在设置内标签排序假设下,该规则在开发标签一致的所有阈值下做出决策。在五个WAM和四个基准设置中,量化产生配置相关的任务损失,这些损失无法仅由位宽或共享偏差阈值解释。在来自两个策略的28个保留配置中,PreDE在观测闭环结果前发布了21个决策(75%覆盖率),所有决策均与观测到的可接受或退化标签匹配。延迟的候选包括可接受结果和33个百分点的损失。在450次Franka Research 3试验中,跨两个独立微调策略,测试前分配到高偏差组的所有配置均表现出显著退化,而低偏差比较未显示统计显著退化。在真实机器人上,W4A4实现了1.37倍的动作查询加速和约44%的峰值内存降低。这些结果支持策略特定的行为校准用于量化配置选择,同时识别需要闭环评估的候选。代码可在该https URL获取。
英文摘要
World action models (WAMs) rely on video-generation backbones, requiring substantial memory and compute for deployment. Post-training quantization reduces memory and can accelerate inference, but bit width, grouping, and quantizer choice define a large configuration space. Identifying configurations that preserve task performance through exhaustive closed-loop evaluation is costly. We propose PreDE (Predict Before You Deploy), a policy-calibrated framework for predicting quantization-induced task degradation from offline action deviations. Using closed-loop outcomes from a small development set, PreDE calibrates two thresholds and accepts, rejects, or defers new configurations using a fixed observation log. Under a within-setting label-ordering hypothesis, the rule issues decisions where all thresholds consistent with the development labels agree. Across five WAMs and four benchmark settings, quantization produces configuration-dependent task losses that cannot be explained by bit width alone or a shared deviation threshold. Across 28 held-out configurations from two policies, PreDE issued 21 decisions before observing closed-loop outcomes (75% coverage), all matching the observed acceptable or degraded labels. Deferred candidates included both acceptable outcomes and a 33-percentage-point loss. In 450 Franka Research 3 trials across two independently fine-tuned policies, all configurations assigned to high-deviation groups before testing showed significant degradation, while low-deviation comparisons showed no statistically significant degradation. On the real robot, W4A4 achieved a 1.37x action-query speedup and approximately 44% lower peak memory. These results support policy-specific behavioral calibration for quantization configuration selection while identifying candidates that require closed-loop evaluation. The code is available at https://github.com/jiuyixu25/PreDE.
Comments9 pages, 3 figures, and 3 tables