Groundbench: 多分辨率多边形定位揭示视觉语言模型中的几何差距
Groundbench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
GroundingBench通过多分辨率多边形定位基准,揭示了视觉语言模型在几何输出上的差距,并指出直接多边形性能远低于边界框,且密集顶点预算下性能崩溃。
中文摘要 AI 辅助
在RefCOCO系列定位任务中,边界框得分已难以区分前沿视觉语言系统,但边界框忽略了物体形状。我们引入了GroundingBench,一个匹配基准,将相同的1500个图像-表达式-指代三元组重新定位到五个顶点预算下的精确N边形。一个固定分母的测试框架分别审计填充区域交并比(IoU)和合法多边形完成度。最强的测试配置在IoU≥0.5(Acc@.5)下达到88.2的框IoU和97.1的准确率,而直接多边形则为57.7和69.2;由于这些主要得分使用不同的参考,我们还将直接多边形与针对相同轮廓目标栅格化的预测框进行比较,在汇总时获得57.7对57.3的结果。性能在N上非单调,并在最密集的预算下崩溃,此时合法性失败加剧了残余几何误差。Qwen的思考设置对比是测试中最大的输入保持配置差异;在冻结模板下,虚假空间线索比虚假颜色线索更具破坏性,且目标偏好可能保持较高而轮廓追踪较差。替代掩码和连续面积评分器保持了主要排序。因此,GroundingBench衡量的是跨越定位、边界构建、序列化和拓扑的操作输出几何差距,而不仅仅是潜在边界感知。
英文摘要
Bounding-box scores on RefCOCO-family grounding leave little room to distinguish frontier vision-language systems, yet boxes discard object shape. We introduce GroundingBench, a matched benchmark that re-targets the same 1,500 image-expression-referent triples to exact-N polygons at five vertex budgets. A fixed-denominator harness separately audits filled-region intersection over union (IoU) and legal-polygon completion. The strongest tested configuration reaches 88.2 box IoU and 97.1 accuracy at IoU >= .5 (Acc@.5), versus 57.7 and 69.2 for direct polygons; because these headline scores use different references, we also compare direct polygons with predicted boxes rasterised against the same contour target, obtaining 57.7 versus 57.3 when pooled. Performance is non-monotone in N and collapses at the densest budget, where legality failures compound residual geometric error. Qwen's thinking-setting contrast is the largest tested input-preserving configuration difference; under frozen templates, false spatial cues are more damaging than false colour cues, and target preference can remain high while contour tracing is poor. Alternate masks and a continuous-area scorer preserve the principal ordering. GroundingBench therefore measures an operational output-geometry gap spanning localisation, boundary construction, serialisation, and topology, rather than latent boundary perception alone.