arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将解耦注意力扩展到多通道图像的密集预测与掩码训练

Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

Umar Marikkar, Sameed Husain, Muhammad Awais, Sara Atito

arXiv 2609.21629首次发表:更新:

发表机构

University of Surrey(萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多通道图像,提出扩展解耦注意力以支持密集预测和掩码训练,通过线性分配恢复跨通道对应,在多个基准上超越最强MC-ViT基线。

AI 中文摘要

多通道成像(MCI)数据与自然图像存在根本性差异,因为每个通道记录的是语义上不同的信号,而非颜色波段。为使视觉编码器适应MCI数据,多通道视觉变换器(MC-ViTs)对每个通道独立进行分词,并将生成的词元连接成一个序列,通道数量不再受架构固定限制。随后,自注意力在所有跨通道的词元上计算,且不限制哪些通道之间可以相互关注,这稀释了各个通道的特征。解耦视觉变换器(DC-ViT)通过将通道内更新与跨通道更新分离,并在通道合并前为每个通道形成独立表示来调节这一问题。然而,其公式按空间位置配对词元,因此要求每个通道具有相同的可见词元。在独立逐通道掩码下的对应关系,通过求解每个通道保留块之间的线性分配来恢复,这使得解耦注意力能够以标准配置而非受限配置与当前的掩码多通道训练相结合。在跨越荧光显微镜、成像质谱流式细胞术和卫星影像的三个分类基准和三个分割基准上(包括高通道数下的密集预测),所提出的公式优于最强的MC-ViT基线。

英文摘要

Multi-Channel imaging (MCI) data differs fundamentally from natural images, as each channel records a semantically distinct signal rather than a colour band. To adapt vision encoders to MCI data, Multi-Channel Vision Transformers (MC-ViTs) tokenize each channel independently and concatenate the resulting tokens into one sequence, and the channel count is no longer fixed by the architecture. Self-attention is then computed across all channel-patch tokens with no restriction on which channels attend to which, which dilutes the features of individual channels. The Decoupled Vision Transformer (DC-ViT) regulates this by separating updates computed within a channel from updates computed across channels, and by forming a representation per channel before the channels are combined. Its formulation, however, pairs tokens by spatial position, and thus requires the same visible tokens in every channel. Correspondence under independent per-channel masking is recovered by solving a linear assignment between the retained patches of each channel, which allows decoupled attention to be combined with current masked multi-channel training in its standard configuration rather than a restricted one. Across three classification and three segmentation benchmarks spanning fluorescence microscopy, imaging mass cytometry and satellite imaging, including dense prediction at high channel counts, the resulting formulation outperforms the strongest MC-ViT baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑