发表机构
Hokkaido University; RIKEN Center for Computational Science (R-CCS); Oak Ridge National Laboratory; A*STAR; Institute of Science Tokyo; Japan Synchrotron Radiation Research Institute (JASRI); RIKEN SPring-8 Center(北海道大学; 理化学研究所计算科学研究中心(R-CCS); 橡树岭国家实验室; 新加坡科技研究局; 东京科学大学; 日本同步辐射研究所(JASRI); 理化学研究所SPring-8中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对十亿像素科学图像,提出结构引导的掩码自编码框架SGMA,结合四叉树分词与结构条件掩码,在多个数据集上显著提升分割性能并加速推理。
AI 中文摘要
自监督预训练与视觉Transformer(包括掩码自编码器,MAE)相结合的方法难以应用于十亿像素级的科学图像。随机掩码与科学数据的结构化、多尺度形态不匹配,而均匀分词会产生过长的序列,使得O(N^2)复杂度的注意力机制变得不切实际。我们提出了SGMA,一种用于超高分辨率科学图像的结构引导掩码自编码框架。SGMA结合了两个组件:一个内容自适应的四叉树分词器,将十亿像素图像压缩为固定长度的序列;以及一个结构条件化的掩码过程,使重建偏向于空间信息丰富的区域。为了在不同尺度上稳定这一过程,我们引入了阻尼累积(DA),它将树中信号相关的响应聚合到一个结构画布上,用于指导掩码。由此产生的预训练任务在保留精细微观结构的同时,与标准ViT编码器和MAE风格的重建保持兼容。在电子显微镜、全切片光学显微镜和X射线CT数据集上,SGMA始终优于MAE基线。它在8K x 8K x 28K的SpringXCT数据集上达到了95.68%的Dice分数,比相同架构的MAE基线提高了+13.00个百分点;在32K^2的WSI PAIP数据集上达到了83.21%的Dice分数,提高了+16.84个百分点,同时提供了高达24.8倍的推理加速。
英文摘要
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.
CommentsAccepted to NeurIPS 2026. 22 pages, 10 figures, 6 tables