arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过强化学习进行推理引导的部件级视觉定位

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning

Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Chia-Wei Tang, Zaber Ibn Abdul Hakim, Anuj Karpatne, Chris Thomas

arXiv 2607.15374首次发表:更新:

AI 中文总结

研究部件级视觉定位难题,提出OP - HRG粗到细推理引导定位策略,先定位父物体再定位部件,经自我检查反思结果,引入部件感知GRPO框架训练,训练的4B模型性能优异且可迁移到推理分割。

AI 中文摘要

多模态大语言模型(MLLMs)能很好地根据自由形式语言查询定位整个物体,但当查询命名的是部件而非物体时就会遇到困难。我们将此归因于缺少物体 - 部件层次结构,因为部件与物体在同一单步中定位。我们提出了物体 - 部件分层反射定位(OP - HRG),这是一种从粗到细的推理引导定位策略,先定位父物体,再定位其中的部件。然后通过自我检查反思结果,并扩展到重新编码预测裁剪区域以检查正在校正的区域。我们引入了一个部件感知GRPO框架,用阶段奖励来训练我们的管道。以这种方式训练的4B模型在PascalPart、PartImageNet和InstructPart上优于7B的定位语言模型和SAM3,并能迁移到推理分割任务中。

英文摘要

Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace this to a missing object-part hierarchy, since parts are localized in the same single step used for objects. We propose Object-Part Hierarchical Reflective Grounding (OP-HRG), a coarse-to-fine reasoning-guided grounding strategy that first localizes the parent object and then the part within it. A self-check then reflects on the result, with an extension to re-encode the predicted crop to inspect the region it is correcting. We introduce a part-aware GRPO framework to train our pipeline with stage-wise rewards. A 4B model trained this way outperforms 7B grounding LLMs and SAM3 across PascalPart, PartImageNet, and InstructPart, and transfers to reasoning segmentation.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑