arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16874cs.CV

用于实例分割的质心位置编码的加速解码

Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation

发表机构摩德纳与雷焦艾米利亚大学
查看机构详情
  • University of Modena and Reggio Emilia(摩德纳与雷焦艾米利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov, Mohammad Mahdi, Yuqian Fu, Giorgia Franchini, Danda Pani Paudel, Marko Bertogna, Luc Van Gool

首次发表
浏览论文内容

中文总结 AI 辅助

针对实例分割中质心位置编码解码开销大的问题,提出优化的CUDA解码算法,显著降低延迟,优于CPU和朴素GPU实现。

中文摘要 AI 辅助

除了模型推理之外,将原始网络输出转换为任务级表示的解码阶段构成了执行成本的很大一部分。尽管其实际影响重大,但预测解码受到的关注相对较少,且通常使用通用的CPU例程或低效的GPU内核来实现,这限制了模型效率提升所带来的好处。在这项工作中,我们研究了与最近用于实例分割的正弦质心编码相关的解码开销,其中每个像素回归其实例质心的位置嵌入。这种方法允许在没有预定义提议的情况下进行灵活分割,但从密集嵌入中提取实例掩码会产生较高的计算成本。我们提出了一种针对该编码量身定制的基于CUDA的优化解码算法实现,明确解决了现代GPU上并行化、同步和内存访问方面的挑战。我们的解决方案显著降低了解码开销并改善了端到端推理延迟,优于基于CPU的方法和朴素的GPU实现。结果表明,高效解码对于充分利用先进输出表示的优势至关重要,并强调了为实时计算机视觉系统联合设计编码方案及其解码算法的重要性。

英文摘要

Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of the execution cost. Despite its practical impact, prediction decoding has received comparatively little attention and is often implemented using generic CPU routines or inefficient GPU kernels, limiting the benefits of advances in model efficiency. In this work, we investigate the decoding overhead associated with a recent sinusoidal centroid encoding for Instance Segmentation, in which each pixel regresses a positional embedding of its instance centroid. This approach allows flexible segmentation without predefined proposals, but extracting instance masks from dense embeddings incurs a high computational cost. We present an optimized CUDA-based implementation of the decoding algorithm tailored to this encoding, explicitly addressing challenges related to parallelization, synchronization, and memory access on modern GPUs. Our solution significantly reduces decoding overhead and improves End-to-End inference latency, outperforming both CPU-based approaches and naive GPU implementations. The results demonstrate that efficient decoding is essential to fully exploit the advantages of advanced output representations and highlight the importance of jointly designing encoding schemes and their decoding algorithms for real-time computer vision systems.

补充信息

↑