arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldDiT:用于世界和动作建模的统一扩散架构

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra

arXiv 2607.23909首次发表:更新:

AI 中文总结

研究提出WorldDiT统一扩散架构,结合动作生成与视觉世界建模,不依赖大型预训练VLM动作主干,在四个LIBERO模拟套件中表现出色,处于帕累托前沿,为未来扩展研究提供了低于十亿参数的基线。

AI 中文摘要

近期许多机器人策略通过使用大型预训练视觉语言模型(VLM)作为动作主干来追求更强控制。我们引入了WorldDiT,一种统一的扩散变压器架构,它将动作生成与视觉世界建模相结合,无需大型预训练VLM动作主干就能实现强大性能。训练时,单个扩散变压器生成连续动作块并从未来相机帧预测归一化RGB补丁目标。在四个LIBERO模拟套件中,WorldDiT在所有报告四个套件的方法中,处于总模型参数和平均成功率的帕累托前沿。这些结果为未来的扩展研究提供了强大的低于十亿参数的基线。

英文摘要

Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.

Comments9 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑