基于新型可供性类型的零样本二维定位
Zero-shot 2D Grounding with Novel Affordance Types
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出新型可供性类型的零样本二维定位任务,构建对应基准,提出AffordAnything及可训练变体AffordAnything+,在AGD20K-NAT基准上较SOTA方法OOAL的IoU@0.4提升12.3%。
AI中文摘要:
二维可供性定位旨在定位人类可与物体交互的区域。现有研究聚焦于识别训练期间所见的可供性类型,未探究模型对新型可供性的泛化能力,而这对实际应用至关重要。我们提出了新型可供性类型(NAT)的零样本二维定位任务,并引入了NAT基准。随后,我们提出了AffordAnything这一无训练方法,其利用分割线索,动机源于可供性区域与物体子部件间的强相关性。为进一步提升性能,我们开发了可训练变体AffordAnything+,该变体学习结合这些线索。在我们提出的AGD20K-NAT基准上,我们的最佳模型AffordAnything+相较于SOTA可供性定位方法OOAL,在IoU@0.4指标上实现了12.3%(绝对)的显著提升。
英文摘要:
2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types seen during training and does not study models' ability to generalize to novel affordances, which is crucial for real-world applications. We propose the task of zero-shot 2D grounding with novel affordance types (NAT) and introduce the NAT benchmarks. We then propose AffordAnything, a training-free method that leverages segmentation cues, motivated by the strong correlation between affordance regions and object subparts. To further improve performance, we develop AffordAnything+, a trainable variant that learns to combine these cues. On the proposed AGD20K-NAT benchmark, our best model AffordAnything+ achieves a substantial improvement of 12.3% (absolute) in IoU@0.4 over the SOTA affordance grounding method, OOAL.