发表机构
Sun Yat-sen University; Hong Kong Polytechnic University(中山大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FAST通过重用单视图基础模型的预训练投影并采用零参数重连策略,构建可扩展的ViT匹配器,在600万对训练数据上实现跨视图密集对应匹配的最先进性能。
AI 中文摘要
扩展已成为语言和视觉基础模型进步的主要驱动力,然而其在精确对应匹配中的作用仍未得到充分探索。在这项工作中,我们提出了Flow Any Scene Transformer(FAST),一种由两个关键见解驱动的可扩展对应模型。首先,我们揭示了单视图视觉基础模型内部的查询-键投影编码了一种粗略但可复用的跨视图匹配先验。其次,在交叉注意力形式中重用这些预训练投影,为基于ViT的匹配器(由单视图编码器构建)提供了高度有效的初始化。在这些见解的指导下,我们在一个普通的单视图基础模型上构建FAST,利用零参数重连策略将选定的自注意力层转换为交叉注意力以进行跨视图交互。这种设计使得基于ViT的匹配器能够随着单视图基础模型的进步而扩展,无需专门的成对中心预训练阶段。为了充分释放这种公式的扩展潜力,我们组装了一个包含600万对训练语料库,用于跨多样共视图像对的通用密集2D位移估计。大量实验表明,FAST在广泛的基准测试中实现了最先进的性能,同时随着骨干网络大小和训练数据的增加而有利地扩展。
英文摘要
Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.