发表机构
Tel Aviv University; Bar-Ilan University(特拉维夫大学; 巴伊兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对实例检索中同一对象表观尺寸不同导致的失效,发现核心原因是O2I比例不匹配,提出查询侧尺度增强与OWLv2裁剪重排序器,在ILIAS 100M上取得最优性能,且O2I鲁棒性可通过LoRA微调学习。
AI 中文摘要
视觉实例检索在查询图像与图库图像中同一对象表观尺寸不同时往往失效。本文表明,主要原因通常不是分辨率损失,而是对象-图像(O2I)比例不匹配:对象在两幅图像中占据的比例不同。在由3021个Objaverse对象在5种相机距离下渲染的受控基准测试中,12种预训练骨干网络中有9种的跨距离性能下降超过80%可归因于O2I不匹配而非分辨率;多尺度架构将纯分辨率影响降至个位数,但仍同样易受O2I不匹配影响。该失效还具有不对称性:紧凑查询图像对宽图库图像的检索可靠性,高于反向情况。基于此分析,查询侧尺度增强与OWLv2裁剪重排序器在ILIAS 100M数据集上达到了最优性能(重排序前mAP@1000为29.2,重排序后为42.0),且无需训练或修改预计算的图库索引;LoRA微调在单次前向传播中即可匹配查询侧增强的增益,表明O2I鲁棒性是可学习的。
英文摘要
Visual instance retrieval often fails when the same object appears at different apparent sizes in the query and gallery. We show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images. On a controlled benchmark of 3,021 Objaverse objects rendered at five camera distances, more than 80% of the cross-distance degradation is attributable to O2I mismatch rather than resolution for 9 of 12 pretrained backbones; multi-scale architectures cut the resolution-only effect to single digits yet remain equally susceptible. The failure is also asymmetric: tight queries retrieve more reliably against wide gallery images than the reverse. Guided by this analysis, query-side scale augmentation and an OWLv2 crop reranker reach state of the art on ILIAS 100M (29.2 mAP@1000 before reranking, 42.0 after) without training or modifying the precomputed gallery index, and a LoRA fine-tune matches the query-side gains at a single forward pass, showing that O2I robustness is learnable.
Comments24 pages. Preprint, under review