arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构引导的掩码自编码器用于超高分辨率科学图像理解

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib

arXiv 2609.30682首次发表:更新:

发表机构

Hokkaido University; RIKEN Center for Computational Science (R-CCS); Oak Ridge National Laboratory; A*STAR; Institute of Science Tokyo; Japan Synchrotron Radiation Research Institute (JASRI); RIKEN SPring-8 Center(北海道大学; 理化学研究所计算科学研究中心(R-CCS); 橡树岭国家实验室; 新加坡科技研究局; 东京科学大学; 日本同步辐射研究所(JASRI); 理化学研究所SPring-8中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对十亿像素科学图像,提出结构引导的掩码自编码框架SGMA,结合四叉树分词与结构条件掩码,在多个数据集上显著提升分割性能并加速推理。

AI 中文摘要

自监督预训练与视觉Transformer(包括掩码自编码器,MAE)相结合的方法难以应用于十亿像素级的科学图像。随机掩码与科学数据的结构化、多尺度形态不匹配,而均匀分词会产生过长的序列,使得O(N^2)复杂度的注意力机制变得不切实际。我们提出了SGMA,一种用于超高分辨率科学图像的结构引导掩码自编码框架。SGMA结合了两个组件:一个内容自适应的四叉树分词器,将十亿像素图像压缩为固定长度的序列;以及一个结构条件化的掩码过程,使重建偏向于空间信息丰富的区域。为了在不同尺度上稳定这一过程,我们引入了阻尼累积(DA),它将树中信号相关的响应聚合到一个结构画布上,用于指导掩码。由此产生的预训练任务在保留精细微观结构的同时,与标准ViT编码器和MAE风格的重建保持兼容。在电子显微镜、全切片光学显微镜和X射线CT数据集上,SGMA始终优于MAE基线。它在8K x 8K x 28K的SpringXCT数据集上达到了95.68%的Dice分数,比相同架构的MAE基线提高了+13.00个百分点;在32K^2的WSI PAIP数据集上达到了83.21%的Dice分数,提高了+16.84个百分点,同时提供了高达24.8倍的推理加速。

英文摘要

Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.

CommentsAccepted to NeurIPS 2026. 22 pages, 10 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑