AI 中文总结
该研究提出主动功能可供性接地任务,构建FUSE框架结合不确定性驱动探索与摊销规划器,基于Habitat基准验证其在非神示接地中性能最优且计算量降低1.33倍。
AI 中文摘要
具身智能体常需基于物体功能而非身份识别并与之交互,这要求它们主动获取能揭示区分性功能证据的观测结果。现有可供性接地方法从固定视角运行,缺乏在功能线索被遮挡或不完整时决定观测位置的机制。我们提出主动功能可供性接地这一新任务,即智能体依次探索场景以识别并空间上接地满足功能查询的物体。为解决该问题,我们提出FUSE,一种自适应语义-几何证据获取框架,其结合显式不确定性驱动探索与学习到的摊销规划器,以高效选择信息丰富的视角。我们还引入基于Habitat的基准用于评估主动功能可供性接地。实验表明,FUSE实现了观测到的非神示接地性能最高,相比完全显式探索减少了1.33倍计算量,且在多种可供性知识源上均保持有效。
英文摘要
Embodied agents must often identify and interact with objects based on their function rather than their identity, requiring them to actively acquire observations that reveal discriminative functional evidence. Existing affordance grounding methods operate from fixed viewpoints and lack mechanisms for deciding where to look when functional cues are occluded or incomplete. We introduce Active Functional Affordance Grounding, a new task in which an agent sequentially explores a scene to identify and spatially ground an object satisfying a functional query. To address this problem, we propose FUSE, an adaptive semantic-geometric evidence acquisition framework that combines explicit uncertainty-driven exploration with a learned amortized planner to efficiently select informative viewpoints. We further introduce a Habitat-based benchmark for evaluating active functional grounding. Experiments show that FUSE achieves the highest observed non-oracle grounding performance while reducing computation by 1.33x relative to fully explicit exploration, and remains effective across multiple affordance knowledge sources.
CommentsUnder review. 15 Pages. 9 tables, 3 Figures