arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AT-ViT:面向标本馆图像植物性状识别的区域定向多视图视觉Transformer,融合交叉注意力与多尺度分块技术

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Prifti

arXiv 2608.21067首次发表:更新:

AI 中文总结

该研究针对标本馆图像植物性状识别的背景干扰问题,提出双分支视觉Transformer模型AT-ViT,通过交叉注意力与掩码引导分块加权机制提升植物区域注意力,在多任务中实现准确率与鲁棒性提升。

AI 中文摘要

从标本馆图像自动识别植物性状对植物科学至关重要,但仍具挑战性,因为背景元素(如文本标签、装裱痕迹、色标)会引入捷径学习,导致模型依赖虚假的非植物线索而非植物形态,这种偏差会降低模型的泛化能力与可解释性。本文提出AT-ViT,这是一种双分支视觉Transformer,通过多尺度、多视图交叉注意力融合方案,联合编码原始标本扫描图像及其经分割得到的对应图像;AT-ViT还引入掩码引导的分块加权机制,放大与植物相关的区域,抑制背景驱动的特征。通过在原始扫描图像上学习,同时受分割掩码通过掩码引导分块重加权机制的引导,该模型被鼓励聚焦于植物器官,更高效地学习以植物为中心的表征。在多个性状分类任务(如叶基形状、刺)中,AT-ViT实现了一致的准确率提升,改进了对植物区域的注意力定位,并在合成背景扰动下表现出更高的鲁棒性。具体而言,相较于CrossViT,AT-ViT大幅提升了空间注意力定位,使植物区域对齐度(平均IoU_p)提高15.66至18.03个百分点,同时降低了背景重叠度(平均IoU_b)27.92至31.02个百分点;在背景噪声条件下,AT-ViT对背景扰动的鲁棒性显著更高,相较于ResNet101的准确率提升最高达32.32个百分点,相较于CrossViT的提升最高达5.07个百分点。

英文摘要

Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models to rely on spurious non-plant cues rather than plant morphology. This bias degrades both generalization and interpretability. In this paper, we introduce AT-ViT, a dual-branch Vision Transformer that jointly encodes raw herbarium scans and their segmented-derived counterparts via a multi-scale, multi-view cross-attention fusion scheme. AT-ViT further incorporates a mask-guided patch weighting mechanism that amplifies plant-relevant regions and attenuates background-driven features. By learning from the original scans while being guided by segmentation masks through the mask-guided patch reweighting mechanism, the model is encouraged to focus on plant organs and learn plant-centric representations more effectively. Across multiple trait classification tasks (e.g., leaf base shape, thorns), AT-ViT delivers consistent accuracy gains, improves attention localization on plant regions, and exhibits increased robustness under synthetic background perturbations. Specifically, AT-ViT substantially improves spatial attention grounding, boosting plant-region alignment (Avg IoU_p: +15.66 to +18.03 pp) while reducing background overlap (Avg IoU_b: -27.92 to -31.02 pp) relative to CrossViT, and remains markedly more robust to background perturbations, outperforming ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points under background-noise conditions.

Journal refIET Computer Vision, 2026, e70059

DOI:10.1049/cvi2.70059

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑