arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MaLViL:用于医学图像分割的多轴低秩视觉长短期记忆网络

MaLViL: Multi-axis Low-rank Vision-LSTM for Medical Image Segmentation

Afshin Bozorgpour, Sina Ghorbani Kolahi, Moein Heidari, Ilker Hacihaliloglu, Dorit Merhof

arXiv 2608.17635首次发表:更新:

发表机构

Faculty of Informatics and Data Science, University of Regensburg; Independent Computer Science Researcher; University of British Columbia; Fraunhofer Institute for Digital Medicine MEVIS(雷根斯堡大学信息学与数据科学学院; 独立计算机科学研究者; 不列颠哥伦比亚大学; 弗劳恩霍夫数字医学梅维斯研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视觉长短期记忆网络(ViL)用于医学图像分割时丢失精细细节、内存开销大的问题,提出MaLViL网络,结合Bi-LRViL、SaLViL、CDM、SGSM等模块,在多个医学图像基准上实现了最优或具竞争力的分割精度,内存降低最多83倍。

AI 中文摘要

视觉长短期记忆网络(ViL)可实现高效的全局建模,但其计算成本仍随空间令牌数量扩展,因此现有分割模型将ViL限制在粗瓶颈层,丢失了精细的解剖细节。将二维特征光栅化为一维序列会进一步破坏正交扫描轴上的邻接关系。我们提出MaLViL,即多轴低秩视觉长短期记忆网络,它将ViL扩展到解码器的所有分辨率。双向低秩ViL(Bi-LRViL)在紧凑的正交子空间中进行推理,并通过正交残差保留细节;感知尺度的SaLViL在序列化前恢复跨轴邻接;跨方向混合器(CDM)融合正交水平和垂直遍历路径;统计引导跳跃调制(SGSM)进一步保留编码器跳跃连接中的边界线索。在皮肤病变、超声和多器官CT基准测试中,MaLViL达到了具有竞争力或最优的分割精度,同时在解码器精细分辨率下将ViL算子的内存降低了多达83倍。代码可在以下URL获取:this https URL。

英文摘要

Vision-LSTM (ViL) enables efficient global modeling, but its cost still scales with the number of spatial tokens, so existing segmenters confine ViL to a coarse bottleneck and lose fine anatomical detail. Rasterizing 2D features into a 1D sequence further breaks adjacency across the orthogonal scan axis. We propose MaLViL, a Multi-axis Low-rank Vision-LSTM network that extends ViL across decoder resolutions. Bidirectional low-rank ViL (Bi-LRViL) reasons on a compact orthonormal subspace and preserves detail through an orthogonal residual; scale-aware SaLViL restores cross-axis neighbors before serialization; and a Cross-Directional Mixer (CDM) fuses orthogonal horizontal and vertical traversal paths. Statistics-Guided Skip Modulation (SGSM) further retains boundary cues in encoder skips. On skin-lesion, ultrasound, and multi-organ CT benchmarks, MaLViL achieves competitive or state-of-the-art segmentation accuracy, while reducing ViL operator memory by up to $83\times$ at fine decoder resolutions. Code is available at: https://github.com/xmindflow/malvil.

CommentsAccepted at the MICCAI Workshop on Machine Learning in Medical Imaging (MLMI), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑