arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PointRL:从可验证标注证据中学习点级视觉-语言接地

PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

Jingyang Su, Pu Cao, Xiuze Jin, Longyue Zhang, Qing Song, Lu Yang

arXiv 2608.25299首次发表:更新:

发表机构

School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(北京邮电大学智能工程与自动化学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PointRL是一种可验证强化学习框架,通过将异构标注证据转换为指向指令并设计多维度奖励,提升了Qwen3.5-4B等模型的点级视觉-语言接地性能,在多个基准上取得显著增益。

AI 中文摘要

视觉-语言模型(VLMs)越来越多地依赖点坐标作为GUI交互、机器人操作和交互式视觉系统中视觉接地的紧凑且可执行接口。然而,学习可靠的指向行为仍然困难,因为监督空间本质上不唯一:同一目标区域内可能有多个有效坐标,而多实例指令需要目标覆盖、计数一致性和重复抑制。本研究提出PointRL,这是一种从现有异构标注证据中学习点级接地的可验证强化学习框架。PointRL将边界框、掩码和实例标签转换为指向指令,同时保留其目标支持、实例成员资格和集合约束作为隐藏验证器证据,即保留在提示外并由确定性检查器用于对预测进行评分的标注。所提出的奖励评估可解析性、点有效性、实例覆盖、基数一致性以及冗余或缺失预测。在PointArena上,PointRL将Qwen3.5-4B的整体准确率从56.11%提升至65.58%。在RoboSpatial、BLINK和Ref-Adv上的进一步评估显示,相同主干模型在评估的外部基准上也获得了提升,表明可验证的点级反馈可能有益于这些场景中的空间接地。

英文摘要

Vision-language models (VLMs) increasingly rely on point coordinates as a compact and executable interface for visual grounding in GUI interaction, robotic manipulation, and interactive visual systems. However, learning reliable pointing behavior remains difficult because the supervision space is inherently non-unique: many coordinates may be valid within the same target region, while multi-instance instructions require target coverage, count consistency, and duplicate suppression. This work presents PointRL, a verifiable reinforcement learning framework that learns point-level grounding from existing heterogeneous annotation evidence. PointRL converts bounding boxes, masks, and instance labels into pointing instructions, while retaining their target supports, instance membership, and set constraints as hidden verifier evidence, i.e., annotations kept outside the prompt and used by a deterministic checker to score predictions. The proposed reward evaluates parseability, point validity, instance coverage, cardinality consistency, and redundant or missing predictions. On PointArena, PointRL improves the overall accuracy of Qwen3.5-4B from 56.11% to 65.58%. Further evaluations on RoboSpatial, BLINK, and Ref-Adv show same-backbone gains on the evaluated external benchmarks, suggesting that verifiable point-level feedback may benefit spatial grounding in these settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑