AI 中文总结
研究部件级视觉定位难题,提出OP - HRG粗到细推理引导定位策略,先定位父物体再定位部件,经自我检查反思结果,引入部件感知GRPO框架训练,训练的4B模型性能优异且可迁移到推理分割。
AI 中文摘要
多模态大语言模型(MLLMs)能很好地根据自由形式语言查询定位整个物体,但当查询命名的是部件而非物体时就会遇到困难。我们将此归因于缺少物体 - 部件层次结构,因为部件与物体在同一单步中定位。我们提出了物体 - 部件分层反射定位(OP - HRG),这是一种从粗到细的推理引导定位策略,先定位父物体,再定位其中的部件。然后通过自我检查反思结果,并扩展到重新编码预测裁剪区域以检查正在校正的区域。我们引入了一个部件感知GRPO框架,用阶段奖励来训练我们的管道。以这种方式训练的4B模型在PascalPart、PartImageNet和InstructPart上优于7B的定位语言模型和SAM3,并能迁移到推理分割任务中。
英文摘要
Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace this to a missing object-part hierarchy, since parts are localized in the same single step used for objects. We propose Object-Part Hierarchical Reflective Grounding (OP-HRG), a coarse-to-fine reasoning-guided grounding strategy that first localizes the parent object and then the part within it. A self-check then reflects on the result, with an extension to re-encode the predicted crop to inspect the region it is correcting. We introduce a part-aware GRPO framework to train our pipeline with stage-wise rewards. A 4B model trained this way outperforms 7B grounding LLMs and SAM3 across PascalPart, PartImageNet, and InstructPart, and transfers to reasoning segmentation.
CommentsECCV 2026