arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36980cs.CV

UltraMatch:用于超快速和内存高效图像匹配的传输路径路由

UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching

Jiajun Le, Yifan Lu, Zizhuo Li, Lei Cao, Junjun Jiang, Jiayi Ma

首次发表
浏览论文内容

中文总结 AI 辅助

UltraMatch通过轻量传输路径路由器仅匹配少量候选路径,实现超快速、内存高效的半稠密图像匹配,速度提升1.67倍,内存仅0.44 GiB,并支持6K分辨率推理。

中文摘要 AI 辅助

尽管在准确性和效率方面取得了近期进展,但在现有的半稠密匹配器中,由于稠密的令牌级匹配,粗匹配仍然是不可或缺且代价高昂的阶段。我们提出了UltraMatch,一个超高效且可扩展的半稠密匹配框架,它通过仅路由一小部分候选匹配路径,绕过了稠密令牌级匹配的二次计算和内存成本。其核心是一个轻量级的传输路径路由器,该路由器在粗块表示上操作,为每个源块对候选目标块进行排序,并仅保留一小部分,从而将后续的令牌级匹配限制在选定的路径上,避免了构建完整的令牌到令牌匹配矩阵。我们进一步设计了一种稀疏的全局双Softmax,它仅在被路由的块候选上进行匹配,同时在稀疏匹配空间中保留全局竞争。除了匹配加速之外,UltraMatch还采用了面向部署的结构重参数化用于特征提取,以及一个具有共享参数的微型精细匹配头,进一步降低了推理成本和内存消耗。UltraMatch在半稠密匹配器中实现了具有竞争力的准确性,同时运行速度比SuperPoint+LightGlue快1.67倍,峰值推理内存仅为0.44 GiB。其可扩展性使得在单个RTX 3090上能够进行高达6K分辨率的推理,而现有的半稠密匹配器在达到2K之前就会耗尽内存。我们的路由策略也是可迁移的,在EDM和ELoFTR中实现了约2倍的端到端加速,且没有精度损失。项目仓库可在该https URL获取。

英文摘要

Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67$\times$ faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2$\times$ end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at https://github.com/JiajunLe/UltraMatch.

发表机构

  • Wuhan University(武汉大学)
  • Xiaomi Corporation(小米公司)
  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑