arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08660cs.CV

CoordFormer:给我任意坐标,我便给你标签

CoordFormer: Give Me Any Coordinates and I Will Give You Labels

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli, Luigi Di Stefano

首次发表
浏览论文内容

中文总结 AI 辅助

CoordFormer提出基于坐标的解码器与局部交叉注意力机制,结合边缘聚焦策略,实现低内存、高精度的超高分辨率图像语义分割,在MaSS13K等基准上达到最优性能。

中文摘要 AI 辅助

在超高分辨率图像上进行语义分割仍然具有挑战性,原因在于高昂的计算成本以及难以捕捉细粒度细节。我们提出了CoordFormer,一种新颖的基于坐标的语义分割架构,它通过配备局部交叉注意力机制的坐标解码器,在任意空间位置预测标签。该解码器将坐标嵌入与高分辨率局部块特征相结合,并与由ViT基础编码器处理的下采样图像提取的全局标记进行交互,从而在保持像素级精度的同时提供丰富的语义上下文。这种设计支持在任意分辨率下进行灵活推理,同时在对超高分辨率输入处理时保持较低的内存占用,并支持一种高效的语义边缘聚焦策略,该策略将计算集中在边界上,在降低延迟和计算成本的同时保持细粒度精度。CoordFormer在MaSS13K上达到了最先进的性能,并在DIS5K和KPIs上优于同等规模及参数更多的现有方法,证明了其在高质量、超高分辨率语义分割中的有效性。

英文摘要

Semantic segmentation on very-high-resolution images remains challenging due to the high computational cost and the difficulty of capturing fine-grained details. We propose CoordFormer, a novel coordinate-based architecture for semantic segmentation that predicts labels at arbitrary spatial locations through a Coordinate Decoder equipped with a Localized Cross-Attention mechanism. The decoder combines coordinate embeddings with high-resolution local patch features and interacts with global tokens extracted from a downsampled image processed by a ViT foundation encoder, enabling rich semantic context while preserving pixel-level precision. This design enables flexible inference at arbitrary resolutions while keeping memory low on very-high-resolution inputs, and supports an efficient semantic-edge-focused strategy that concentrates computation along boundaries, maintaining fine-grained accuracy while reducing latency and computational cost. CoordFormer achieves state-of-the-art performance on MaSS13K and outperforms comparably sized and higher-parameter methods on DIS5K and KPIs, demonstrating its effectiveness for high-quality, very-high-resolution semantic segmentation.

发表机构

  • CVLab, University of Bologna(博洛尼亚大学CV实验室)
  • Ca’ Foscari University of Venice(威尼斯大学)
  • SINA, company(SINA公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑