VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding
专题命中 空间理解 :point cloud(abstract);分类 cs.CV
AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。