发表机构
American International University Bangladesh; United International University; ELITE Research Lab(孟加拉国美国国际大学; 联合国际大学; 精英研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对孟加拉国淡水鱼类识别,审计了CLIP等多款零样本视觉-语言模型,发现BioCLIP2性能最优,模型得分受生物专业化、命名等多因素影响。
AI 中文摘要
零样本视觉-语言模型(VLMs)正越来越多地被用作无需训练的物种识别工具,但报告的准确率所反映的信息不止于视觉物种知识。我们针对来自孟加拉国两个来源的7种淡水鱼类类别(共10321张图像),对CLIP、BioCLIP、BioCLIP2及多语言Jina CLIP v2对照组开展审计。BioCLIP2在BFF-15数据集(采用英文通用名称)上的准确率达72.36%,在SylFishBD数据集(采用学名)上为68.91%,而通用CLIP的对应准确率仅为25.15%和14.40%。BioCLIP2的孟加拉语提示在平衡准确率上接近随机水平(14.22-14.29%);Jina CLIP则将孟加拉语的区分度部分恢复至21.89%和16.36%,但纯孟加拉语名称在两个数据集上的准确率回落至14.29%。配对的SylFishBD干预实验显示,弱模糊无显著影响,强模糊/灰度掩码会造成适度损失,白色掩码伪影更明显,且存在强物种依赖性。因此,零样本生物VLM的得分同时反映了生物专业化程度、多语言对齐情况、命名规则、提示表述及上下文因素。
英文摘要
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.