arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越未来预测:去噪作为机器人控制的生成式自适应

Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control

Zanyi Wang, Yuheng Lei, Dengyang Jiang, Ping Luo, Mengdi Wang, Zhixuan Liang, Shilong Liu

arXiv 2609.28339首次发表:更新:

发表机构

HKU; HKUST; Princeton University; UC San Diego(香港大学; 香港科技大学; 普林斯顿大学; 加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出NowWAM,通过去噪当前观测并预测动作,无需未来目标,在LIBERO-Plus上以更少视觉令牌和更快速度超越基线,验证了去噪轨迹作为控制接口的有效性。

AI 中文摘要

预训练的生成式扩散变换器(DiTs)通过大规模图像和视频生成训练,捕获了丰富的像素级视觉和语言条件结构。越来越多的机器人策略构建在这一生成先验之上,但如何将其迁移到控制任务中仍不明确,现有方法通常通过未来视觉预测来实现这种迁移。我们提出了一个更基本的问题:预训练的生成式DiT实际上对动作学习贡献了什么,以及这一先验应如何为控制任务进行自适应。我们引入了NowWAM,一种无需未来目标的协同训练框架,该框架对当前观测进行去噪,并从同一视觉流中预测机器人动作,直接将原生生成目标与跨去噪轨迹的动作导向表征耦合。在匹配的受控设置下,过去和未来的视觉目标表现相当,而将训练限制在干净端点会大幅降低鲁棒性,这表明单独的未来目标对于生成式自适应并非必需,而去噪轨迹仍然是控制的有效接口。在LIBERO-Plus上,NowWAM使用FLUX2-Klein达到87.7%的准确率,比未来目标协同训练基线提高了6.1个百分点,同时将训练视觉令牌减半(从784降至392),并将步进时间从2.85秒缩短至1.63秒,实现了1.8倍的加速。使用纯文本到图像的Z-Image骨干网络,NowWAM进一步达到87.8%的准确率,表明强大的控制自适应并不依赖于视频生成或图像编辑骨干网络。

英文摘要

Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.

CommentsProject page: https://xmz111.github.io/NowWAM/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑