arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23832cs.CV

DINOv2在活动性沙眼分类中的比较性能与参数高效适配

Comparative Performance and Parameter-Efficient Adaptation of DINOv2 for Active Trachoma Classification

Kibrom Gebremedhin, Hadush Hailu, Bruk Gebregziabher, Yordanos Hailu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在统一协议下比较了六个预训练骨干网络,发现DINOv2结合高效通道注意力(仅5个可学习参数)在沙眼分类中达到91.66%准确率,提供了参数高效的分类基准。

中文摘要 AI 辅助

结膜照片的自动分级可以降低沙眼患病率调查的成本和变异性,但在统一协议下,现代预训练视觉表征、轻量级特征适配和训练目标设计的相对价值尚未确立。本研究利用来自公共UCSF/Lietman数据集的1,546张图像,对沙眼性炎症-滤泡(TF)与正常进行二分类的受控评估。图像使用OPTED流程进行零样本睑结膜分割、对齐、裁剪和标准化。我们首先使用通用分类流程比较六个预训练骨干网络,然后在DINOv2 ViT-B/14上评估四种轻量级适配机制。在分层五折交叉验证下,采用高效通道注意力(ECA)和焦点加中心损失的DINOv2达到了91.66±0.97%的准确率、90.69±1.10%的宏F1分数和96.06±0.71%的AUC。ECA仅引入五个可学习参数,同时匹配了显著更大替代方案的性能。目标消融进一步表明,ECA在不同损失函数下并未一致地改进普通DINOv2;最低方差91.66%的准确率是通过交叉熵加中心损失获得的。总体而言,微调的DINOv2表征提供了大部分预测性能,而ECA提供了一种高度参数高效的细化,其效果取决于训练目标。所得工作流程为活动性沙眼图像分类提供了可复现的基准。

英文摘要

Automated grading of conjunctival photographs could reduce the cost and variability of trachoma prevalence surveys, but the relative value of modern pretrained visual representations, lightweight feature adaptation, and training-objective design has not been established under a common protocol. This study presents a controlled evaluation for binary classification of Trachomatous Inflammation-Follicular (TF) versus Normal using 1,546 images from the public UCSF/Lietman collection. Images are processed using the OPTED pipeline for zero-shot tarsal-conjunctiva segmentation, alignment, cropping, and standardization. We first compare six pretrained backbones using a common classification pipeline and then evaluate four lightweight adaptation mechanisms on DINOv2 ViT-B/14. Under stratified five-fold cross-validation, DINOv2 with Efficient Channel Attention (ECA) and focal-plus-center loss achieved 91.66 +/- 0.97% accuracy, 90.69 +/- 1.10% macro-F1, and 96.06 +/- 0.71% AUC. ECA introduces only five learnable parameters while matching the performance of substantially larger alternatives. Objective ablation further showed that ECA did not consistently improve plain DINOv2 across loss functions; the lowest-variance 91.66% accuracy was obtained with cross-entropy plus center loss. Overall, the fine-tuned DINOv2 representation provided most of the predictive performance, while ECA offered a highly parameter-efficient refinement whose effect depended on the training objective. The resulting workflow provides a reproducible benchmark for active trachoma image classification.

发表机构

  • Mekelle University(默克莱大学)
  • Maharishi International University(玛赫西国际大学)
  • Signal Technologies(信号技术公司)
  • MicroLink Information Technology College(微联信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑