arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.30484cs.RO

ELAN4D:以具身为中心的4D监督用于视觉-语言-动作模型的即插即用适配

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation

  • Torr Vision Group, University of Oxford(托尔视觉组,牛津大学)
  • The Chinese University of Hong Kong, Shenzhen(香港大学(深圳))
  • Tsinghua University(清华大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • University College London(伦敦大学学院)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

Zeyuan He, Bowen Yang, Zhirui Fang, Keru Zhou, Lei Jiang, Jingjing Qian, Fan Mo, Junchi Yan, Philip Torr, Xiu Li, Li Jiang, Jialin Yu

更新

AI总结:

提出ELAN4D框架,通过未来机器人关键点轨迹作为预测性时空监督,以即插即用方式增强VLA策略的鲁棒性和泛化能力。

AI中文摘要:

视觉-语言-动作(VLA)模型在机器人操作中展现出潜力,但现有大多数策略通过直接从当前观测回归动作来反应式运行,没有显式建模未来动态。这限制了它们在分布外扰动下的泛化能力。为解决此问题,我们提出ELAN4D,一个以具身为中心的4D感知训练框架,通过未来机器人关键点轨迹作为预测性时空监督来增强VLA策略。仅利用本体感觉状态的前向运动学,我们推导出机器人关键点(如关节和末端执行器)的3D位移轨迹,预处理成本可忽略。这些轨迹提供度量且紧凑的监督,无需外部跟踪器或重建。一个即插即用的辅助分支,配备轻量级轨迹解码器,在通过梯度隔离保护预训练视觉-语言主干的同时,将4D信号注入动作专家。推理时丢弃轨迹解码器,保持基础策略接口不变。在LIBERO、LIBERO-Plus、RoboTwin2.0和真实世界操作任务上的大量实验表明,ELAN4D持续优于强VLA基线,在分布外扰动(包括相机、背景和布局变化)下取得最佳整体性能和显著提升。这些结果凸显了以具身为中心的4D监督对于构建更鲁棒和可泛化的操作策略的有效性。

英文摘要:

Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observations, without explicitly modeling future dynamics. This limits their ability to generalize under out-of-distribution perturbations. To address this issue, we propose ELAN4D, an embodiment-centric, 4D-aware training framework that enhances VLA policies with future robot keypoint tracks as predictive spatio-temporal supervision. Using only forward kinematics from proprioceptive states, we derive 3D displacement tracks of robot keypoints, such as joints and the end-effector, with negligible preprocess cost. These tracks provide metric and compact supervision without requiring external trackers or reconstruction. A plug-and-play auxiliary branch with a lightweight track decoder injects this 4D signal into the action expert while preserving the pretrained vision-language backbone through gradient isolation. The track decoder is discarded during inference, leaving the base policy interface unchanged. Extensive experiments on LIBERO, LIBERO-Plus, RoboTwin2.0 and real-world manipulation tasks demonstrate that ELAN4D consistently improves over strong VLA baselines, achieving the best overall performance and substantial gains under out-of-distribution perturbations, including camera, background, and layout shifts. These results highlight the effectiveness of embodiment-centric 4D supervision for building more robust and generalizable manipulation policies.

↑