AI 中文总结
本研究提出基于深度可分离分解的参数高效三维 U-Net,在肝脏分割任务中以 24 倍参数缩减超越稠密卷积,并证明空间滤波器放置比总量更重要。
AI 中文摘要
三维稠密卷积网络在体素医学图像分割中表现最强,但其参数数量扩展性差:将稠密 k×k 卷积扩展到 k×k×k 会使权重增加 k 倍。我们观察到深度可分离分解不存在这一代价。由于立方核项仅应用于深度卷积阶段,而主导参数数量的逐点投影保持不变,相同架构从二维扩展到三维时参数仅增加 5%,而稠密卷积 U-Net 则增加 200%。我们利用这一不对称性构建了一个具有 536,990 个参数的三维 U-Net,比相同结构的稠密三维 U-Net 参数少 24 倍。在 MSD Task03 肝脏任务(医学分割十项全能肝脏任务,源自 LiTS)上,对所有 131 个公开体素进行五折交叉验证逐例评估,模型达到肿瘤 Dice 0.577(95% 置信区间 [0.518, 0.633])和肝脏 Dice 0.947(95% 置信区间 [0.941, 0.952])。其肝脏 Dice 超过先前报告的性能。其肿瘤 Dice 比低分辨率配置(0.4701)高 0.107,比二维配置(0.5394)高,而参数约为其二十四分之一,面内分辨率约为其一半。在相同条件下于公共留出分割上训练,其在肝脏上超过稠密三维 U-Net +0.031 Dice(配对 p = 0.006),在肿瘤上超过 +0.041(95% 置信区间 [+0.005, +0.081],配对 p = 0.056),表明该分解在医学成像特有的小数据场景中起到正则化作用。我们进一步在 LiTS 和二维内窥镜基准上证明,网络中大部分可学习的空间滤波器可以被固定位移替代而不损失精度,但全部替代则明显变差,空间容量的放置比其总量更重要。
英文摘要
Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share this penalty. Because the cubic kernel term applies only to the depthwise stage while the pointwise projection, which dominates the parameter count, is unchanged, the same architecture grows by 5 % from 2D to 3D where a dense convolutional U-Net grows by 200 %. We exploit this asymmetry to build a 3D U-Net with 536,990 parameters, 24x fewer than an identical dense 3D U-Net. On MSD Task03 Liver (the Medical Segmentation Decathlon liver task, derived from LiTS), evaluated per case under five-fold cross-validation over all 131 public volumes, the model reaches a tumor Dice of 0.577 (95% CI [0.518, 0.633]) and a liver Dice of 0.947 (95% CI [0.941, 0.952]). Its liver Dice exceeds previously reported performance. Its tumor Dice exceeds their low-resolution configuration (0.4701) by 0.107 and their 2D configuration (0.5394), at approximately one twenty-fourth of the parameters and roughly half the in-plane resolution. Trained under identical conditions on a common held-out split, it exceeds a dense 3D U-Net on liver by +0.031 Dice (paired p = 0.006) and on tumor by +0.041 (95 % CI [+0.005, +0.081], paired p = 0.056), suggesting the factorization also acts as a regularizer in the small-data regime characteristic of medical imaging. We further show, on both LiTS and a 2D endoscopy benchmark, that a large fraction of the network's learnable spatial filters can be replaced by fixed shifts at no cost in accuracy, but that replacing all of them is measurably worse, the placement of spatial capacity matters more than its total amount.