AI 中文总结
该研究提出Hi-Token分层坐标分词方法,结合基于几何的Hi-GAR奖励,在生成式视觉定位任务中提升了多模型多基准的定位性能,优于强专用基线。
AI 中文摘要
生成式视觉语言模型(VLMs)通常将边界框坐标视为独立的输出符号,使得数值顺序和轴语义隐含。我们发现这种表示是视觉定位中错误的重要来源。Hi-Token用针对百位、十位和个位的轴特定标记对每个坐标进行编码,在保留现有VLM架构的同时,增加了由粗到细的结构并提高了标记复用率。Hi-GAR通过基于几何的奖励补充这种表示,用于组相对策略优化(GRPO),使用多尺度的框重叠和坐标精度。在匹配训练条件下的受控比较显示,Hi-Token在整个评估的IoU范围内提高了定位性能,Hi-GAR进一步减少了低重叠预测且仅在训练期间使用。在三个VLM骨干和RefCOCO系列上的实验显示,在模型和基准上均获得一致提升,Hi-R1在大多数报告的指标上比强专用基线实现了更高值。对标记频率、数字边界、对象尺度和IoU分布的分析解释了坐标表示和奖励训练的效果,结果表明结构化坐标生成是生成式视觉定位的有效方法。
英文摘要
Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify this representation as an important source of error in visual grounding. Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM architecture. Hi-GAR complements this representation with a geometry-based reward for Group Relative Policy Optimization (GRPO), using box overlap and coordinate accuracy at multiple scales. Controlled comparisons under matched training conditions show that Hi-Token improves localization throughout the evaluated IoU range. Hi-GAR further reduces low-overlap predictions and is used only during training. Experiments on three VLM backbones and the RefCOCO family show consistent gains across models and benchmarks. Hi-R1 achieves higher values than strong specialist baselines on most reported metrics. Analyses of token frequency, digit boundaries, object scale, and IoU distributions explain the effects of coordinate representation and reward training. The results show that structured coordinate generation provides an effective approach to generative visual grounding. Project page: https://xyzzzh.github.io/Hi-Token/
Comments15 pages, 7 figures, 15 tables