arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19553cs.CV

IoU曲线上定位精度的核心所在:无标注推理时的边界优化

Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary Refinement

Bo Ma

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无标注推理时的边界优化(LFPR)方法,在多个视觉-语言定位数据集上提升了IoU曲线各精度指标,证明指代选择与边界精度部分可分离,为定位精度优化提供了新思路。

中文摘要 AI 辅助

视觉-语言模型能够识别出正确的指代对象,但返回的边界框不够精确。我们研究在推理阶段,一个冻结的直接回答模型是否可以利用自身预测结果,在不访问目标标注的情况下分配一次额外的局部观测。无标注精度优化(LFPR)将预测的小区域送入更高分辨率的通道,在上下文裁剪区域内重新定位表达式,仅在固定几何约束下接受候选框,并返回固定的坐标中点。我们在三个证据层级上报告了结果:在31921个回顾性Ref-L4表达式上,LFPR将mAcc$_{0.5:0.95}$从72.947%提升至76.013%(Acc@0.5从88.531%提升至89.725%,Acc@0.9从55.788%提升至61.142%);在30969个RefCOCO/RefCOCO+/RefCOCOg表达式上的冻结迁移,在Acc@0.5、mAcc和平均IoU上均提升了各数据集的性能(合并mAcc提升0.645,Acc@0.5提升0.817),而Acc@0.9整体不变:仅路由操作在此指标上提升1.162个点,但裁剪、约束和融合操作使其下降1.192,形成抵消而非无严格IoU效应;在图像不重叠的前瞻性Flickr30K Entities评估中,所有端点均得到提升(mAcc提升0.973,Acc@0.9提升1.022),单框变体下提升更显著(mAcc提升2.575,Acc@0.9提升3.689);将相同操作应用于两个已发布的定位专家模型,在延迟约翻倍的情况下,所有端点均得到提升(EGM-4B/8B的Acc@0.9分别提升1.569/6.716),与专家训练结合而非替代它;移除相同候选框约束的无约束对照组在所有指标上均弱于现有模型,表明约束是关键组成部分。综合结果显示,指代对象选择与边界精度部分可分离,不同组件会在IoU曲线上移动相反区域,这是单一阈值无法揭示的行为。

英文摘要

Vision--language models can identify the correct referent while returning an imprecise bounding box. We study whether a frozen direct-answer model can use its own prediction to allocate one additional localized observation without accessing target annotations at inference. Label-free precision refinement (LFPR) routes predicted-small regions to a higher-resolution pass, re-grounds the expression inside a context crop, admits a candidate only under fixed geometric guards, and returns a fixed coordinate-wise midpoint. We report results across three evidence tiers. On 31,921 retrospective Ref-L4 expressions, LFPR raises mAcc$_{0.5:0.95}$ from 72.947\% to 76.013\% (Acc@0.5 88.531\%$\to$89.725\%, Acc@0.9 55.788\%$\to$61.142\%). A frozen transfer to 30,969 RefCOCO/RefCOCO+/RefCOCOg expressions improves every dataset at Acc@0.5, mAcc, and mean IoU (pooled mAcc $+0.645$, Acc@0.5 $+0.817$), while Acc@0.9 is unchanged overall: routing alone gains $+1.162$ points there, but crop, guards, and fusion give back $-1.192$, offsetting rather than showing no strict-IoU effect. A prospective, image-disjoint Flickr30K Entities evaluation improves every endpoint (mAcc $+0.973$, Acc@0.9 $+1.022$), more strongly under a single-box variant (mAcc $+2.575$, Acc@0.9 $+3.689$). The same operator applied to two released grounding specialists improves every endpoint (Acc@0.9 $+1.569$/$+6.716$ for EGM-4B/8B) at roughly twice the latency, composing with specialist training rather than replacing it. A genuine unguarded control (guard removed from the same candidates) underperforms the incumbent on every metric, showing the guard is load-bearing. Together, these results show that referent selection and boundary precision are partially separable, with different components moving opposing regions of the IoU curve -- behavior a single threshold cannot reveal.

发表机构

  • Auckland University of Technology(奥克兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑