发表机构
Plano West Senior High School; Centennial High School; The University of Texas at Dallas(普莱诺西高中; 百年高中; 德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估边缘可部署的2-8B视觉语言模型在物种识别上的表现,发现所有模型在野外图像上均显著退化,且专用模型BioCLIP虽小却大幅领先,表明差距源于训练数据而非模型规模。
AI 中文摘要
相机陷阱通常在边缘硬件上于野外运行,这些硬件可能只有有限或没有网络连接,这使得小型、可本地部署的视觉语言模型(VLM)——而非前沿规模的模型——成为评估物种识别任务时实际相关的类别。我们测试了处于此部署相关2-8B参数范围内的模型是否携带真正的分类学知识,在96个物种的任务上评估了四个此类VLM(Qwen3-VL 2B/4B/8B,Gemma3 4B),并与领域专用专家模型BioCLIP(300M参数)进行对比,比较了干净的iNaturalist照片与来自6个此http URL数据集的相机陷阱图像,在两组独立采样的评估集上进行。所有模型识别物种的表现远高于随机水平,但每个模型——无论是通用型还是专用型——在野外图像上性能都急剧下降(领域差距为9.6至26.6个百分点,在不同分类学层级和两个评估集上均一致),这表明性能下降反映的是图像整体可读性问题,而非细粒度判别失败。尽管BioCLIP的规模远小于所有测试的VLM,但其性能大幅优于每个VLM(在扩展的200图像样本上,每个模型均高出33.2至59.2个百分点),这表明差距源于专用训练数据而非模型规模;然而,BioCLIP自身的领域差距(18.0个百分点)在统计上与最佳VLM的领域差距(22.3个百分点)无法区分,这表明从干净图像到野外图像的性能下降本身是图像质量变化的一个属性,而非通用模型的弱点。在开放集提示下,5.9%至9.6%的响应是语法有效但分类学上不存在的物种名称;各模型间相对捏造率的排序在两个评估集上完全一致,这比任何单一的点估计都更为稳健。
英文摘要
Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against the domain-specific specialist BioCLIP (300M parameters) on a 96-species task, comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.science collections, on two independently-sampled evaluation sets. All models identify species far above chance, but every model -- general-purpose or specialist -- degrades sharply on field imagery (domain gaps of 9.6--26.6 percentage points, consistent across taxonomic levels and both evaluation sets), indicating the degradation reflects general image legibility rather than fine-grained discrimination failure. BioCLIP substantially outperforms every VLM tested (by 33.2--59.2 percentage points across an expanded 200-image sample for every model) despite its far smaller size, suggesting the gap reflects specialized training data rather than model scale; yet BioCLIP's own domain gap (18.0 points) is statistically indistinguishable from the best VLM's (22.3 points), suggesting the clean-to-field degradation itself is a property of the image-quality shift rather than a general-purpose-model weakness. Under open-set prompting, 5.9--9.6% of responses are syntactically valid but taxonomically nonexistent species names; the relative fabrication-rate ranking across models replicates exactly across both evaluation sets, a more robust finding than any single point estimate.