发表机构
Taiyuan University of Technology(太原理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GSToken,一种带显式几何信息的高斯令牌,通过冻结令牌生成器的评估协议验证其在多模态脑肿瘤分割中,比容量匹配基线更优,可提升三维医学图像表示的信息密度。
AI 中文摘要
多模态MRI的有效分割是提高神经网络在脑肿瘤识别中准确率的核心。现有方法通常通过固定补丁编码或学习注意力池化(如TokenLearner)将三维体积压缩为令牌序列,但这些压缩方案会丢弃显式空间形状信息,生成的令牌无法传递病变形态或空间范围的概念。同时,端到端评估会将令牌生成器的信息保留能力与下游解码器的重建能力纠缠在一起,且不同方法间缺乏统一的容量约束,导致性能差异难以归因。本文首次将高斯令牌引入多模态脑肿瘤分割:每个令牌不仅携带语义特征,还包含学习到的三维中心、各向异性尺度和方向,以可忽略的参数成本为表示赋予显式几何支撑。我们进一步提出冻结令牌的效用评估协议:冻结训练后的令牌生成器,将其输出转换为固定容量的序列化约束,并使用共享的轻量Transformer探针在严格匹配条件下独立测量每个令牌生成器保留的信息。多配对统计测试显示,在冻结探针条件下,GSToken在容量匹配的自适应基线中始终大幅优于后者,在所有肿瘤亚区、表面及距离指标上均表现出一致优势。这些结果表明,在令牌中显式编码空间几何可显著提高体积表示的信息密度,为紧凑三维医学图像表示及下游读取提供了新的设计原则。
英文摘要
Effective segmentation of multi-modal MRI is central to improving neural network accuracy in brain tumor recognition. Existing methods typically compress 3D volumes into token sequences via fixed patch encoding or learned attention pooling (e.g., TokenLearner). However, these compression schemes discard explicit spatial shape information; the resulting tokens convey no notion of lesion morphology or spatial extent. Meanwhile, end-to-end evaluation entangles a tokenizer's information retention with the reconstruction capacity of the downstream decoder, and the lack of a unified capacity contract across methods makes performance differences difficult to attribute. In this paper, we introduce Gaussian tokens to multi-modal brain tumor segmentation for the first time: each token carries not only a semantic feature but also a learned 3D center, anisotropic scale, and orientation, endowing the representation with explicit geometric support at negligible parameter cost. We further propose a frozen-token utility evaluation protocol: the trained tokenizer is frozen, its output is cast into a fixed-capacity serialized contract, and a shared lightweight Transformer probe independently measures each tokenizer's retained information under strictly matched conditions. Multi-seed paired statistical testing shows that GSToken consistently and substantially outperforms capacity-matched adaptive baselines under frozen probing, with uniform advantages across all tumor sub-regions, surface, and distance metrics. These results demonstrate that explicitly encoding spatial geometry within tokens significantly improves the information density of volumetric representations, offering a new design principle for compact 3D medical image representation and downstream reading.