学习频谱分配:用于自适应体积分割的分数扩散框架
Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation
浏览论文内容
中文总结 AI 辅助
提出FHEAT-Seg,通过分数扩散算子让优化器自适应决定各层频谱混合,实现轻量级3D分割,在半监督和全监督下均超越基线并大幅降低计算量。
中文摘要 AI 辅助
我们解决了3D医学图像分割中的自适应计算问题:与其设计另一个骨干网络,我们询问每个网络阶段需要多少频谱混合,并让优化器来回答。我们从分数热方程的离散余弦变换(DCT)解中推导出FHEAT,一个双参数算子族。分数阶alpha和扩散强度D控制该算子,当D=0时它恰好是恒等算子。通过半群时间tau = D*alpha重新参数化,相同分辨率的实例精确组合,因此跨相同分辨率阶段的任何扩散分布相当于一个学习强度的单一Sobolev型正则化器。这种恒等极限让每层的优化器(而非设计者)决定是否需要全局混合以及其锐度。我们将FHEAT实例化在一个轻量级U形架构(Light-UNETR)中,并配以具有自适应有理激活的Kolmogorov-Arnold混合器(KAN3D),产生FHEAT-Seg。在三个公共基准上以5%到20%的标签率,训练产生梯度驱动的频谱稀疏化:八个阶段级算子中有七个将D驱动到零,幸存者在提供半监督注意力图的解码器层中饱和于最锐利的低通(alpha约0.9)。被淘汰的层在推理时成为精确的恒等捷径,将FLOPs从4.29G降至0.90G(下降79%),参数为0.975M。在标准半监督协议下,FHEAT-Seg达到Dice分数90.47%(左心房)、78.79%(胰腺CT)和81.90%(BraTS 2019),领先于五种半监督方法和Light-UNETR基线。大变体在全监督下也超过Light-UNETR-L(Dice 93.09%、85.11%和87.19%),参数为2.851M,FLOPs为55.75G。这些结果表明频谱计算的分配是优化动态的可学习属性,而非手动设计承诺。
英文摘要
We address adaptive computation in 3D medical image segmentation: instead of designing another backbone, we ask how much spectral mixing each network stage needs and let optimization answer. We derive FHEAT, a two-parameter operator family, from the discrete cosine transform (DCT) solution of a fractional heat equation. A fractional order alpha and a diffusion strength D govern the operator, and at D=0 it is exactly the identity. Reparametrized by the semigroup time tau = D*alpha, same-resolution instances compose exactly, so any distribution of diffusion across same-resolution stages amounts to a single Sobolev-type regularizer of learned strength. This identity limit lets the optimizer of each layer, not the designer, decide whether global mixing is needed and how sharp it should be. We instantiate FHEAT in a lightweight U-shaped architecture (Light-UNETR) paired with a Kolmogorov-Arnold mixer (KAN3D) with adaptive rational activations, yielding FHEAT-Seg. At 5% to 20% label rates on three public benchmarks, training produces gradient-driven spectral sparsification: seven of the eight stage-level operators drive D to zero, and the survivor saturates at the sharpest low-pass (alpha ~ 0.9) in the decoder layer feeding the semi-supervised attention map. The retired layers become exact identity shortcuts at inference, cutting FLOPs from 4.29G to 0.90G (a 79% drop) at 0.975M parameters. Under a standard semi-supervised protocol, FHEAT-Seg reaches Dice scores of 90.47% (left atrium), 78.79% (Pancreas-CT), and 81.90% (BraTS 2019), ahead of five semi-supervised methods and the Light-UNETR baseline. The large variant also surpasses Light-UNETR-L under full supervision (Dice 93.09%, 85.11%, and 87.19%) with 2.851M parameters and 55.75G FLOPs. These results suggest that the allocation of spectral computation is a learnable property of optimization dynamics, not a manual design commitment.
发表机构
- Fujian Medical University(福建医科大学)
- Stockholm University(斯德哥尔摩大学)
机构由 AI 辅助整理,请以论文原文为准。