TorchMorph:CUDA加速的形态学变换
TorchMorph: CUDA-accelerated Morphological Transforms
查看机构详情
- Shanghai University(上海大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出轻量PyTorch扩展TorchMorph,实现22个CUDA加速形态学算子,吞吐量远超CPU参考,可便捷迁移现有流程,用于GPU视觉任务。
中文摘要 AI 辅助
形态学变换是用于形状与掩码处理的经典工具,但Python生态中事实上的参考实现scipy.ndimage.morphology仅支持CPU、单数组操作,需进行代价高昂的设备-主机往返传输,无法在GPU训练循环中使用。基于PyTorch的GPU视觉库仅覆盖了这些算子的窄小子集,通常限于二维空间和扁平结构元素。本文提出TorchMorph,一款填补该空白的轻量PyTorch扩展,提供22个公共算子,涵盖二值形态学、灰度形态学、精确与近似距离变换、熵正则化最优传输,全部实现为融合CUDA内核,可直接处理最多8个空间维度的(B, C, Spatial...)格式CUDA张量。其API刻意与scipy.ndimage.morphology保持参数完全一致,包括边界模式、结构元素原点与预分配输出,只需修改导入语句即可迁移现有流程。本文描述了分层架构与各算子族背后的内核设计。与单线程CPU参考实现相比,灰度形态学的批量执行吞吐量可达scipy.ndimage.morphology的1100倍,精确欧氏距离变换可达350倍,Sinkhorn求解器比POT快42倍;二值与倒角算子可精确复现SciPy对应结果,所有浮点算子与CPU参考的绝对误差均在1.8e-6以内。TorchMorph以MIT许可发布,代码见指定链接。
英文摘要
Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e. scipy.ndimage, is CPU-only, single-array, and therefore unusable inside a GPU training loop without an expensive device-to-host round trip. GPU vision libraries built on PyTorch cover a narrow subset of these operators, typically restricted to two spatial dimensions and flat structuring elements. We present TorchMorph, a lightweight PyTorch extension that closes this gap. TorchMorph exposes 22 public operators covering binary morphology, greyscale morphology, exact and approximate distance transforms, and entropy-regularised optimal transport, all implemented as fused CUDA kernels that operate directly on (B, C, Spatial...) CUDA tensors with up to eight spatial dimensions. The API deliberately mirrors scipy.ndimage argument-for-argument, including border modes, structuring-element origins and pre-allocated outputs, so that existing pipelines port with a change of import. We describe the layered architecture and the kernel designs behind each operator family. Against single-threaded CPU references, batched execution reaches up to 1.1e3 times the throughput of scipy.ndimage on greyscale morphology and up to 350x on exact Euclidean distance transforms, while the Sinkhorn solver runs up to 42x faster than POT. Binary and chamfer operators reproduce their SciPy counterparts exactly, and every float-valued operator agrees with the CPU reference to within 1.8e-6 absolute error. TorchMorph is released under the MIT licence at https://intcomp.github.io/tm.