A4A:从人类示范中跨实体迁移面向动作的4D可操作性
A4A: Cross-Embodiment Transfer of Action-Oriented 4D Affordances from Human Demonstrations
浏览论文内容
中文总结 AI 辅助
提出面向动作的4D可操作性表示,通过预测交互点轨迹预训练机器人策略,实现从人类示范到机器人的跨实体操作知识迁移,实验验证其提升VLA策略性能。
中文摘要 AI 辅助
人类示范中包含丰富的操作知识,但目前尚不清楚哪些信息可以有效地迁移到机器人控制中。现有的可操作性表示通常被表述为2D掩码、3D区域、接触点或可执行性评分,因此主要识别交互可能发生的位置。然而,有效的操作还需要建模交互相关几何形状在任务执行过程中如何演变。为弥合这一差距,我们引入了面向动作的4D可操作性,它表示交互相关3D点的语言条件未来轨迹。这些轨迹捕获任务条件下的几何演变,而非特定实体的动作,从而实现了跨人类和机器人的可迁移交互先验。基于这一表示,我们从现有的人-物交互视频数据和互补的RGB-D示范中构建了一个大规模面向动作的4D可操作性数据集,并引入了A4A,一个可操作性到动作的框架,该框架在操作微调之前使用4D可操作性轨迹预测来预训练机器人策略。在仿真和真实世界中的实验验证了A4A的有效性,表明使用面向动作的4D可操作性数据进行预训练持续提高了多种VLA策略的操作性能。这些结果确立了面向动作的4D可操作性作为一种有效的跨实体表示,用于将操作知识从人类示范迁移到机器人控制。
英文摘要
Human demonstrations contain rich manipulation knowledge, but it remains unclear what information can be transferred effectively to robot control. Existing affordance representations are typically formulated as 2D masks, 3D regions, contact points, or actionability scores, and therefore primarily identify where interaction may occur. However, effective manipulation also requires modeling how interaction-relevant geometry evolves during task execution. To bridge this gap, we introduce action-oriented 4D affordances, which represent the language-conditioned future trajectories of interaction-relevant 3D points. These trajectories capture task-conditioned geometric evolution rather than embodiment-specific actions, enabling transferable interaction priors across humans and robots. Based on this representation, we construct a large-scale action-oriented 4D affordance dataset from existing human--object interaction video data and complementary RGB-D demonstrations, and introduce A4A, an affordance-to-action framework that uses 4D affordance trajectory prediction to pretrain robot policies before manipulation fine-tuning. Experiments in both simulation and the real world validate the effectiveness of A4A, showing that pretraining with action-oriented 4D affordance data consistently improves the manipulation performance of diverse VLA policies. These results establish action-oriented 4D affordances as an effective cross-embodiment representation for transferring manipulation knowledge from human demonstrations to robot control.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Rutgers University–New Brunswick(罗格斯大学新布朗斯维克分校)
- Nanyang Technological University(南洋理工大学)
- The Hong Kong University of Science and Technology (GZ)(香港科技大学(广州))
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。