发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对植被样方图像多物种植物识别难题,基于微调的DINOv2 ViT-L/14分类器,采用多尺度切片分解、kNN检索、栖息地适应降级等方法,在PlantCLEF 2026挑战赛中获第三名,私有排行榜宏F1为0.43902。
AI 中文摘要
本文描述了DS@GT ARC在植被样方图像多物种植物识别的PlantCLEF 2026挑战赛中获得第三名的解决方案。系统要在仅使用单标签植物图像训练的情况下,预测高分辨率样方照片中的所有物种。该管道围绕微调后的DINOv2 ViT-L/14分类器构建,对每个样方进行多尺度切片分解,切片预测与FAISS kNN检索器融合,并通过源感知时间融合、栖息地适应降级和地理掩码后处理。消融实验表明栖息地适应降级和多尺度聚合贡献最大。一些训练方向未取得成果,推理时的实例感知分割裁剪也未提升性能。所选提交在私有排行榜上的宏F1为0.43902(第三名;公开为0.51096)。
英文摘要
This paper describes DS@GT ARC's third-place solution to the PlantCLEF 2026 challenge on multi-species plant identification in vegetation quadrat images, where systems must predict every species present in high-resolution (~3000 x 3000 pixel) plot photographs while training only on single-label images of individual plants. The pipeline is built around a fine-tuned DINOv2 ViT-L/14 classifier applied over a multi-scale tile decomposition of each quadrat, with per-tile predictions blended with a FAISS kNN retriever and post-processed by source-aware temporal fusion across repeated plot visits, a habitat-fit demotion that injects geographic and altitude priors from the training data, and a South-Western Europe geographic mask. Habitat-fit demotion and multi-scale aggregation are the largest individual contributors in the ablations. Two complementary training-centric directions, a cross-region transformer with noisy-student distillation on the LUCAS dataset and a label-as-query transformer decoder over synthetic CLS-domain pseudo-quadrats, yielded null results. An inference-time augmentation with instance-aware segmentation crops also did not improve performance. The selected submission reaches a private-leaderboard macro-F1 of 0.43902 (third place; public 0.51096); an unselected configuration of the same pipeline scored above 0.45 on the private set. Code: https://github.com/dsgt-arc/plantclef-2026.