arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10981cs.CV

ThinkAfford:面向杂乱场景中细粒度3D grounding的以可供性为中心的推理

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

Xinrui Lin, Sha Zhang, Shumin Wang, Zenghuan Zhu, Jiajun Deng, Yanyong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ThinkAfford模型,通过解耦可供性提议生成与指令推理,结合GRPO优化,在SceneFun3D验证集上的3D grounding指标优于基线方法。

中文摘要 AI 辅助

任务驱动的3D可供性grounding旨在在杂乱的3D场景中定位能实现自然语言指令指定动作的功能区域。现有方法要么直接预测3D掩码,要么通过选择和融合中间2D/3D区域构建掩码,但仍易受两种相互关联的失败模式影响:预测或选定的区域可能遗漏目标交互区域或粒度不合适,而语言grounding在关系指令下可能混淆视觉相似的备选对象。为此,我们引入ThinkAfford,它将高召回率的可供性提议生成与基于指令的推理解耦。具体而言,可供性提议生成(APG)模块首先使用可学习的可供性提示和多级视觉特征预测交互条件热图,提取数量可变的细粒度提议,无需解析对象或部件名称作为分割提示。视觉提示可供性推理(VPAR)随后使用完整指令对标记的提议覆盖区域进行推理,以结构化的“思考后回答”响应返回标识符。此外,分组相对策略优化(GRPO)使用来自提升的3D重叠的提议级奖励,使VPAR选择与最终3D grounding对齐。在SceneFun3D验证集上,ThinkAfford在官方评估器下达到10.69%的AP50和25.46%的AP25,优于可比的3D开放词汇及基于视觉语言模型的2D转3D基线。模块级诊断进一步显示,APG在25%交并比下达到77.5%的召回率,而GRPO训练的VPAR在APG覆盖的查询上达到72.1%的选择准确率,相比之下,监督微调下的准确率为63.4%。

英文摘要

Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing methods either predict 3D masks directly or construct them by selecting and fusing intermediate 2D/3D regions. However, they remain vulnerable to two intertwined failure modes: the predicted or selected regions may miss the target interaction area or have unsuitable granularity, while language grounding may confuse visually similar alternatives under relational instructions. To this end, we introduce ThinkAfford, which decouples high-recall affordance proposal generation from instruction-grounded reasoning. Specifically, the Affordance Proposal Generation module first uses learnable affordance prompts and multi-level visual features to predict interaction-conditioned heatmaps, extracting a variable number of fine-grained proposals without parsed object or part names as segmentation prompts. Visual-Prompted Affordance Reasoning then reasons over labeled proposal overlays using the full instruction, returning identifiers in a structured "think-then-answer" response. Moreover, Group Relative Policy Optimization uses proposal-level rewards from lifted 3D overlap to align VPAR selection with final 3D grounding. On the SceneFun3D validation split, ThinkAfford achieves 10.69% AP50 and 25.46% AP25 under the official evaluator, outperforming comparable 3D open-vocabulary and vision-language-model-based 2D-to-3D baselines. Module-level diagnostics further show that APG attains 77.5% recall at 25% intersection-over-union, while GRPO-trained VPAR achieves 72.1% selection accuracy on APG-covered queries, compared with 63.4% under supervised fine-tuning.

补充信息

↑