arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05215cs.ROcs.CV

VLAff:用于统一可操作 affordance 的视觉-语言-affordance 模型

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

Jihoon Oh, Kento Kawaharazuka, Kei Okada

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出 VLAff 模型,结合 EgoAffordance 数据集,解决人类与机器人 embodiment 不匹配问题,实现视觉 affordance 预测及真实机器人零样本操作等应用。

中文摘要 AI 辅助

从人类视频中学习操作技能是可扩展机器人学习的可行方向,但人类与机器人之间的 embodiment 不匹配使其颇具挑战。一种有前景的解决方案是学习与 embodiment 无关的、以物体为中心的可操作 affordance。本研究提出一种框架,利用以自我为中心的人类视频,结合最先进的 3D 运动恢复结构(Structure-from-Motion)和手部网格重建技术,提取视觉、抓取、轨迹等可操作 affordance,这些 affordance 明确编码了交互位置、抓取方式和移动方式。我们构建了 EgoAffordance 大规模数据集,包含 20.4 万集、560 万个视觉 affordance 以及 1160 万个抓取和轨迹 affordance。在此基础上,我们推出基于大型视觉-语言模型的统一基础模型 VLAff,该模型学习所有可操作 affordance 之间的跨模态关联。给定视觉观测和指令,VLAff 生成视觉 affordance 热力图、抓取位姿和轨迹,随后利用 3D 场景信息将其转换为可直接执行的动作。通过大量实验,我们证明 VLAff 不仅在视觉 affordance 预测上达到最先进性能,还可有效应用于真实机器人任务,如零样本操作和 affordance 引导的机器人学习。

英文摘要

Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionable affordances that are embodiment-agnostic. In this work, we propose a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move. We construct EgoAffordance, a large-scale dataset comprising 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances. Building on this, we introduce VLAff, a large vision-language model-based unified foundation model that learns cross-modal correlations across all actionable affordances. Given a visual observation and instruction, VLAff generates visual affordance heatmaps, grasp poses, and trajectories, which are then converted into directly executable actions by utilizing 3D scene information. Through extensive experiments, we demonstrate that VLAff not only achieves state-of-the-art performance on visual affordance prediction, but can also be effectively applied to real robot applications such as zero-shot manipulation and affordance-guided robot learning.

发表机构

  • University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑