发表机构
The University of Hong Kong; The Hong Kong University of Science and Technology (Guangzhou); Peking University; The Chinese University of Hong Kong; National University of Singapore; Knowin AI(香港大学; 香港科技大学(广州); 北京大学; 香港中文大学; 新加坡国立大学; 考拉悠然科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AffordanceWAM,通过可供性感知的生成式世界动作模型,联合预测未来视觉与动作,利用人类视频实现无需动作标签的机器人操作迁移,并在多个基准上取得一致提升。
AI 中文摘要
可泛化的机器人操作需要预测场景将如何演变、识别何处交互可行,并确定如何行动。带有动作标注的机器人视频直接监督控制,但成本高昂且多样性有限,而第一视角人类视频捕捉了多样的交互,但缺乏机器人动作,且在具身形态和外观上存在差异。我们提出了AffordanceWAM,一种可供性感知的生成式世界动作模型,它在生成的未来世界中通过标量可供性和可供性热图来表示以物体为中心的时空可供性。这种表示将视觉预测锚定在与任务相关的物体和交互区域上以生成动作,并在人类和机器人视频之间提供共享的交互目标。基于预训练的视频扩散Transformer,AffordanceWAM使用分别参数化的世界专家和动作专家,通过掩码联合自注意力耦合,在统一的流匹配目标下联合预测未来RGB观测、标量可供性场、可供性热图和连续机器人动作。人类视频监督所有三个未来世界流,而机器人轨迹额外提供动作监督,从而无需人类动作标签或重定向即可实现迁移。在RoboCasa、CALVIN ABC→D和真实世界操作上的实验表明,与仅使用RGB和仅使用机器人数据的基线相比,性能持续提升。在固定的机器人监督下,随着可供性标注的人类视频规模增加,RoboCasa性能单调提升。这些结果支持可供性作为视觉-语言-动作学习和人-机器人迁移的有效接口。
英文摘要
Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act. Action-labeled robot videos directly supervise control but are costly and limited in diversity, whereas egocentric human videos capture diverse interactions but lack robot actions and differ in embodiment and appearance. We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap, within the generated future World. This representation grounds visual prediction in task-relevant objects and interaction regions for action generation, and provides shared interaction targets across human and robot videos. Built on a pretrained video diffusion Transformer, AffordanceWAM uses separately parameterized World and Action Experts, coupled through Masked Joint Self-Attention, to jointly predict future RGB observations, Scalar Affordance fields, Affordance Heatmaps, and continuous robot actions under a unified flow-matching objective. Human videos supervise all three future-World streams, whereas robot trajectories additionally provide action supervision, enabling transfer without human action labels or retargeting. Experiments on RoboCasa, CALVIN ABC$\rightarrow$D, and real-world manipulation demonstrate consistent gains over RGB-only and robot-data-only baselines. Under fixed robot supervision, RoboCasa performance improves monotonically as affordance-annotated human video scales. These results support affordance as an effective interface for both vision-language-action learning and human-to-robot transfer.