arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29460cs.CVcs.AI

AgriCountDINO:农业中参数高效的示例引导计数与定位

AgriCountDINO: Parameter-Efficient Exemplar-Guided Counting and Localization in Agriculture

  • University of Bologna(博洛尼亚大学)
  • Politecnico di Milano(米兰理工大学)
  • Forschungszentrum Jülich(于利希研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Shengjie Guo, Xin Li, Borjana Arsova, Hanno Scharr, Silvio Salvi

AI总结:

AgriCountDINO提出参数高效的示例引导计数与定位框架,利用冻结DINOv3特征和少量参数,在TPC-268上实现低误差计数并提升零样本泛化能力。

AI中文摘要:

植物及其器官的准确计数和定位支持表型分析和产量估算,然而目标的外观、尺度和密度在不同物种和成像条件下差异很大。示例框指定目标而无需针对特定类别进行重新训练,点预测则识别构成计数的各个实例。我们提出了AgriCountDINO,一个参数高效的示例引导框架,用于联合计数和定位。它根据示例的外观和大小调节冻结的多尺度DINOv3特征,然后逐步将其解码为目标点。漏检目标恢复将监督扩展到初始匹配遗漏的目标,而示例自适应点NMS根据示例尺度过滤重复预测。AgriCountDINO仅需840万个可训练参数,约为TasselNetV4的十分之一,在TPC-268基准上实现了三次样本的MAE为11.92,将计数误差降低了9.7%,同时提供了单个目标的位置。仅在TPC-268上训练,它在FSC-147中未见过的通用目标类别上实现了零样本MAE为14.25,比最佳对比零样本方法提高了6.0%,而无需目标域训练或微调。

英文摘要:

Accurate counting and localization of plants and their organs support phenotyping and yield estimation, yet target appearance, scale, and density vary widely across species and imaging conditions. Exemplar boxes specify the target without category-specific retraining, and point predictions identify the individual instances contributing to the count. We introduce AgriCountDINO, a parameter-efficient exemplar-guided framework for joint counting and localization. It conditions frozen multiscale DINOv3 features on exemplar appearance and size, then progressively decodes them into target points. Missed-object recovery extends supervision to targets overlooked by initial matching, and exemplar-adaptive point NMS filters duplicate predictions according to exemplar scale. With 8.4M trainable parameters, approximately one-tenth of TasselNetV4's, AgriCountDINO achieves a three-shot MAE of 11.92 on the TPC-268 benchmark, reducing counting error by 9.7\% while providing individual target locations. Trained only on TPC-268, it achieves a zero-shot MAE of 14.25 on unseen generic object categories in FSC-147, improving upon the best compared zero-shot method by 6.0\% without target-domain training or fine-tuning.

↑