AffordAny:基于单目RGB图像的开放世界3D affordance grounding,通过视觉-语言引导的几何推理实现
AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning
浏览论文内容
中文总结 AI 辅助
AffordAny是端到端框架,通过单目RGB图像结合VLM引导解码器与伪标签自训练,实现开放世界3D affordance grounding,在未见对象、类别等评估中表现出有效性与鲁棒性。
中文摘要 AI 辅助
开放世界3D affordance grounding需要根据自由形式的语言查询在三维空间中定位功能对象部件。现有方法通常假设存在预先构建的以对象为中心的三维几何结构和封闭的 affordance 本体,这限制了从原始RGB观测数据进行部署。我们提出AffordAny,这是一种端到端框架,它利用单目RGB图像构建大规模文本条件三维部件监督,通过冻结的视觉-语言模型(VLM)引导的解码器进行 affordance 定位,并通过伪标签自训练提升开放世界泛化能力。我们的自动化流程生成了包含5334个对象和10633个部件级样本的基准数据集,涵盖473个类别,类别多样性较现有工作提升了一个数量级。该解码器通过空间投影、指令条件语义压缩以及几何-语义双向交互,逐步融合冻结的Cosmos-2B特征与三维几何。最小扰动伪标签自训练进一步在无需人工标注的情况下新增对象。在评估未见对象、未见类别和未见指令 paraphrases 的系统泛化协议下,我们的方法在自训练后对未见对象实现0.428 IoU,对未见类别实现0.315 IoU,未见类别平均IoU相对提升6.3%(p<0.01),且指令敏感度差距仅为0.105,证明了我们方法的有效性和鲁棒性。
英文摘要
Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D geometry and closed affordance ontologies, limiting deployment from raw RGB observations. We present AffordAny, an end-to-end framework that uses one monocular RGB image to construct large-scale text-conditioned 3D part supervision, ground affordances with a frozen vision-language model (VLM) guided decoder, and improve open-world generalization through pseudo-label self-training. Our automated pipeline produces a benchmark of 5,334 objects and 10,633 part-level samples spanning 473 categories, an order-of-magnitude increase in categorical diversity over prior work. The decoder progressively fuses frozen Cosmos-2B features with 3D geometry through spatial projection, instruction-conditioned semantic compression, and bidirectional geometry-semantics interaction. Minimal-perturbation pseudo-label self-training further adds new objects without human annotation. Under a systematic generalization protocol evaluating unseen objects, unseen categories, and unseen instruction paraphrases, our approach achieves 0.428 IoU on unseen objects and 0.315 IoU on unseen categories after self-training, with unseen-category mIoU improving by 6.3% relative (p<0.01) and an instruction sensitivity gap of only 0.105, demonstrating effectiveness and robustness of our method.