ImCorr:通过隐式特征解码实现亚像素语义对应
ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding
- Pukyong National University(釜庆国立大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
ImCorr通过隐式特征解码在连续特征场上进行亚像素语义对应,消除了网格量化误差,在SPair-71k和AP-10K上显著提升了细粒度匹配精度。
中文摘要 AI 辅助
现代语义对应方法在标准阈值下表现强劲,但在细粒度阈值下性能急剧趋于平稳。我们认为,这种平稳并非源于骨干网络特征的表征能力,而是源于网格绑定的读取方式。基于补丁的视觉变换器将图像标记化到离散网格上,引入了两种形式的量化误差:在源侧,查询最近补丁特征而非精确关键点;在目标侧,缺乏表示精确地面真值位置的网格特征。我们在SPair-71k数据集的全部499,188个关键点上量化了这一量化上限:在标准448x448、补丁大小为14的设置下,84.9%的地面真值关键点在PCK@0.01下没有网格特征表示其精确位置。这是表征层面的结构性限制,与匹配策略无关。我们通过ImCorr(通过隐式特征解码实现亚像素语义对应)解决这一问题,该方法在可于任意连续坐标查询的连续特征场上进行对应估计。一个受FiLM条件调节的解码器被训练用于将亚像素位置信息嵌入特征场。直接在精确关键点坐标处查询该场,理论上消除了源侧表征层面的量化误差,而解码到比骨干网格更密集的网格则显著减少了目标侧的量化误差。在SPair-71k和AP-10K(种内、跨种和跨科)上,ImCorr在细粒度阈值(PCK@0.01-0.05)下提升了性能,在SPair-71k的PCK@0.01上比先前最先进方法提升了6.2个百分点。这些结果表明,表征连续性是为精确语义对应提供有效解决方案。代码可在以下网址获取:此https URL。
英文摘要
The strong performance that modern semantic correspondence methods achieve at standard thresholds plateaus sharply at fine-grained thresholds. We argue that this plateau stems not from the representational capacity of backbone features, but from a grid-tied readout. Patch-based vision transformers tokenize images onto discrete grids, introducing two forms of quantization error: querying nearest patch features instead of exact keypoints on the source side, and the absence of grid features representing precise ground-truth locations on the target side. We quantify this quantization ceiling across all 499,188 keypoints in SPair-71k: under the standard 448x448, patch-14 setting, 84.9% of ground-truth keypoints have no grid feature representing their precise location at PCK@0.01. This is a structural limitation at the representation level, independent of the matching strategy. We address this with ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding, which formulates correspondence estimation over a continuous feature field queryable at arbitrary continuous coordinates. A FiLM-conditioned decoder is trained to embed sub-pixel positional information into the feature field. Querying the field directly at exact keypoint coordinates theoretically eliminates representation-level quantization error on the source side, while decoding onto a grid denser than the backbone grid substantially reduces quantization error on the target side. On SPair-71k and AP-10K (intra-species, cross-species, and cross-family), ImCorr improves performance at fine-grained thresholds (PCK@0.01-0.05), achieving a 6.2 percentage point gain over the prior state of the art at PCK@0.01 on SPair-71k. These results demonstrate that representational continuity is an effective solution for precise semantic correspondence. Code is available at https://github.com/YusungChoi/ImCorr.