从未来隐状态推断机器人操纵动作
DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
查看机构详情
- DeepLeap Research(DeepLeap研究院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出无需视频生成的机器人操纵模型DELE-w0.5,通过从捕获动作相关物理结果的未来隐状态推断动作,在4项长程操纵任务的480次试验中,其性能优于最强基线47.5和30.7个百分点,实现最优表现。
中文摘要 AI 辅助
世界动作模型(World-Action Models, WAMs)基于视频生成主干构建机器人控制,可联合预测密集的未来视觉轨迹与机器人动作。我们认为视频生成是世界动作建模中不必要的中间目标。对于机器人操纵而言,世界模型的目标并非复现世界在每个中间时刻的外观,而是预测执行动作后世界将达到的状态;中间帧仅描述物理状态间的视觉过渡,会消耗大量模型容量与计算资源,却未直接指定机器人动作预期达成的物理结果。本文提出DELE-w0.5,该模型无需依赖视频生成,即可从预测的未来状态推断机器人动作序列,具体而言,它从对应紧凑的未来隐状态中推断动作序列,该未来隐状态捕获机器人交互中与动作相关的物理结果,是连接世界建模与动作生成的明确桥梁。DELE-w0.5的核心设计原则是建模物理世界在机器人动作下的变化,而非其外观逐帧演变的过程,该公式消除了密集视频表示引入的高维视觉冗余,因此支持更廉价的训练与低延迟推理。在四个长程操纵任务的480次真实机器人试验中,我们的DELE-w0.5在所有对比策略中取得最优性能,整体全任务成功率达62.5%,宏观有序阶段进展达81.3%,分别超越最强基线47.5和30.7个百分点。
英文摘要
World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 640 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5% overall full-task success and 81.3% macro ordered-stage progress. It outperforms the strongest baseline by 32.5 percentage points in full-task success and 20.1 percentage points in macro progress.