arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dynin-Robotics:全能统一扩散视觉-语言-动作模型

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do

arXiv 2609.13053首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Dynin-Robotics,基于全能掩码扩散骨干,通过共享轨迹模型统一视觉目标与动力学预测,实现动作生成与选择,提升机器人操作性能。

AI 中文摘要

视觉目标和动力学预测可以为语言条件机器人策略提供目标结果和动作相关场景变化的表示。我们通过共享轨迹模型将这些预测引入动作生成和选择。Dynin-Robotics在Dynin-Omni(一个全能掩码扩散骨干网络)上实现了这一公式,将语言、视觉观察、目标和动作表示为离散令牌。通过改变条件和目标跨度,同一模型学习动作预测、动作条件下一观察预测、终端目标状态预测和轨迹到指令重建。这些接口通过目标预测、动作候选评估以及动作和未来状态预测的联合细化支持测试时扩展。我们在来自48个Open X-Embodiment数据集的大约133万条轨迹上持续预训练模型,并分别适应下游领域。在两个VLABench任务上,机器人预训练在固定的Stage-2步骤预算内提高了适应性,完整的客观混合在相同耦合解码器下比仅策略后训练提高了转移指令成功率。将目标引导与联合动作-下一状态去噪相结合,进一步提高了转移指令成功率,优于仅动作解码;其益处取决于预测的组合方式。Dynin-Robotics在LIBERO和零样本LIBERO-Plus上取得了有竞争力的性能,并在Franka Research 3机器人上四个操作条件下平均成功率达到78.4%。优化的块并行实现相对于报告的分析设置下的基础实现,将模型侧动作解码加速了最多29.2倍。这些结果支持共享轨迹建模作为学习互补机器人目标并在控制期间组合其预测的通用接口。

英文摘要

Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.

Comments36 pages, 13 figures, 15 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑