发表机构
Southern University of Science and Technology; The University of Hong Kong; Shenzhen Loop Area Institute(南方科技大学; 香港大学; 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AnyStep-WAM框架,通过预算对齐蒸馏和自适应推理,在保持任务成功率的同时大幅减少世界动作模型的去噪步数,提升单步生成性能。
AI 中文摘要
世界动作模型(WAMs)将预测性视觉建模与动作生成相结合,通常依赖固定去噪步数的迭代去噪过程。然而,操作任务中的动作块对生成误差的敏感度各不相同:关键动作需要高精度,而敏感度较低的动作允许以更少的去噪步数实现更快生成。为此,我们提出AnyStep世界动作模型(AnyStep World Action Model),这是一个支持可调预算预测和场景相关计算分配的通用的框架。我们的预算对齐教师轨迹蒸馏方法利用显式的冻结教师转换和共享低秩适配器训练区间条件流映射,支持从单步预测到多步细化的动作生成。基于此能力,一个轻量级风险收益调度器从单次单步预览中预测基于教师曲率的难度和特定预算的师生保真度,选择预测满足风险自适应保真度要求的最小预算。我们在三个广泛使用的WAM模型Motus、FastWAM和LingBotVA上,使用RoboTwin 2.0数据集评估了我们的框架。我们的方法分别将平均去噪步数减少了60.2%、49.8%和85.28%,同时保持了基线任务成功率。特别是,我们的AnyStep训练显著提升了模型在单步去噪预算下的性能,在Motus、FastWAM和LingBotVA上的任务成功率分别提高了7.07%、12.08%和8.94%。在六个真实世界操作任务上的实验进一步验证了其有效性。
英文摘要
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.