arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉Transformer对小麦表型分析的合并容忍度有多高?

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

Simon Ravé, Pejman Rasti, David Rousseau

arXiv 2608.23142首次发表:更新:

发表机构

LARIS, University of Angers; UMR INRAE-IRHS(昂热大学LARIS研究所; INRAE-IRHS联合研究单位)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对小麦表型分析任务,系统基准测试了ToMe和互相对合并的合并容忍度,发现分类对合并容忍度高,检测和分割受多种因素限制,部署价值需在目标运行时分析。

AI 中文摘要

基于视觉的小麦表型分析需要在部署约束下进行重复测量,涵盖从生长阶段识别到麦穗计数和器官分割的任务。普通视觉Transformer(ViT)为这些任务提供了通用架构,但二次注意力限制了高通量和边缘推理。无需训练的令牌合并技术颇具吸引力,因为它可插入已训练好的模型而无需重新训练。我们针对生长阶段分类、麦穗检测和小麦器官分割任务,对ToMe(令牌合并)和互相对合并进行了系统基准测试,测量了任务质量、吞吐量、令牌数量和峰值GPU内存,还补充了树莓派5(Raspberry Pi 5)的测量数据。该基准测试揭示了明确的层级:分类对合并的容忍度很高,而检测和分割受重复实例、细器官、密集边界、重建及运行时开销的限制。优化后的注意力后端可消除明显的加速效果,因此部署价值必须在目标运行时上进行分析,而非通过令牌数量推断。

英文摘要

Vision-based wheat phenotyping requires repeated measurements under deployment constraints, from growth-stage recognition to wheat-head counting and organ segmentation. Plain Vision Transformers (ViTs) provide a common architecture for these tasks, but quadratic attention limits high-throughput and edge inference. Training-free token merging is attractive because it can be inserted into trained models without retraining. We provide a systematic benchmark of ToMe and Mutual Pair Merging across growth-stage classification, wheat-head detection, and wheat-organ segmentation, measuring task quality, throughput, token count, and peak GPU memory, with additional Raspberry Pi 5 measurements. The benchmark reveals a clear hierarchy: classification is highly merge-tolerant, while detection and segmentation are constrained by repeated instances, thin organs, dense boundaries, reconstruction, and runtime overhead. Optimized attention backends can erase apparent speedups, so deployment value must be profiled on the target runtime rather than inferred from token count.

CommentsAccepted to the CVPPA workshop at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑