Bridge-WA: 预测世界变化的位置和方式以支持机器人动作
Bridge-WA: Learning Action-Relevant World Dynamics for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
提出Bridge-WA轻量框架,通过蒸馏未来变化教师模型为三个紧凑先验,引导动作生成聚焦于场景变化的位置和方式,提升机器人操作的成功率和鲁棒性。
中文摘要 AI 辅助
通用视觉-语言-动作模型受益于大型视觉-语言先验,但有效操作还需要预测与动作相关的场景变化。现有的世界-动作模型通常依赖大型生成式世界模型或密集的未来展开,这些方法成本高昂,并将容量浪费在与控制弱耦合的视觉细节上。我们提出Bridge-WA,一个轻量级世界-动作框架,将冻结的未来变化教师模型蒸馏为三个紧凑先验:用于预期结果的未来令牌、用于干预支持的变化图,以及用于局部过渡方向的运动流图。WorldBridge通过多源注意力记忆和时空偏置将这些先验条件化到动作变换器上,而在推理时移除教师模型。在VLABench、RoboTwin2.0、LIBERO-Plus和真实机器人评估中,Bridge-WA提高了任务成功率、进展和鲁棒性,在分布外视觉偏移下尤其明显。通过将动作生成聚焦于场景变化的位置和方式,Bridge-WA抑制了背景、光照和干扰物等无关外观因素,从而在不进行部署时密集未来图像生成的情况下实现更好的泛化。代码和可视化可在以下网址获取:this https URL。
英文摘要
General-purpose VLA models leverage large-scale vision-language priors to understand scenes and instructions, but primarily generate actions directly from the current observation. WAMs further model future scene states, offering a broader view of how the environment may evolve. However, for robotic manipulation, predicting the entire future scene can introduce information beyond what is necessary for action generation; more importantly, effective actions require knowing not only what the scene may become, but also where and how relevant changes unfold. To bridge action generation with these action-relevant aspects of the world, we present Bridge-WA, a general world-action framework that learns complementary representations of future states, spatial changes, and local motion. Specifically, Bridge-WA consists of a Latent World Dynamics Module (LWDM) and WorldBridge. LWDM predicts future states, spatial changes, and local motion from VLM outputs, supervised by corresponding world targets. The WorldBridge are embedded into the action transformer and inject layer-specific combinations of these world priors through multi-source attention, spatiotemporal biases, and reliability-gated feature modulation. This design grounds action generation in action-relevant future dynamics while adaptively regulating world guidance, enabling robust generalization to visual disturbances and viewpoint shifts. We evaluate Bridge-WA on four simulation benchmarks and the real robots, where it achieves maximum success-rate improvements of 11.1%, 42.0%, 3.7%, 23.4%, and 11.1%, respectively, and achieves state-of-the-art average success rates on LIBERO-Dynamic, RoboTwin 2.0 and real-world. In particular, Bridge-WA demonstrates strong generalization to visual variations and viewpoint shifts in both simulation and real-world settings. Code and visualizations are available at: https://hcplab-sysu.github.io/BRIDGE-WA.
发表机构
- Sun Yat-sen University(中山大学)
- Pengcheng Laboratory(鹏城实验室)
- Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
- X-Era AI Lab(X-Era AI实验室)
机构由 AI 辅助整理,请以论文原文为准。