发表机构
University of California, Davis(加州大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建农业基准,发现VLM视觉编码器已具备足够特征,通过基于评分标准的验证和概率枢轴锦标赛,显著提升分类性能,揭示模型知识未充分展现。
AI 中文摘要
视觉语言模型(VLMs)在农业分类方面显示出潜力,但在疾病、害虫、损害、质量和物种识别上的零样本性能仍然较差,且尚不清楚这是否反映了视觉特征薄弱,还是未能将这些特征与领域知识联系起来。我们构建了一个包含116个数据集、834个类别和8,324张图像的基准,涵盖这些任务,以隔离差距产生的环节。线性探测表明,VLM视觉编码器已经编码了农业特征,其可分性几乎与自监督DINOv3基线相当,从而排除了弱视觉表示作为主要瓶颈的可能性。将每个模型置于一个参考描述(其参数知识的上界)的条件下,几乎弥合了无辅助下界所留下的大部分差距,表明VLMs对农业的了解比它们展现的更多。为了在推理时无需参考描述的情况下弥合这一差距,我们将测试时推理结构化为围绕固定的、每任务诊断评分标准:模型生成K个候选响应,一个概率枢轴锦标赛(PPT)验证器,根据评分标准进行成对评分,选择最佳响应。这使评判F1得分几乎翻倍,超过下界,并在多个任务上达到或超过上界,尤其将Gemma 4 E4B-it的疾病F1得分提升至0.71,高于其自身0.60的上界。然而,验证器的字母等级置信度得分产生了与预期相反的效果:过滤到其最自信的预测并未提高准确性,并且在所有测试的模型和池大小中与正确性呈负相关,因此该得分不能作为预测不确定性的度量,观察到的收益大部分可能来自基于评分标准的生成,而非成对验证。
英文摘要
Vision-language models (VLMs) show promise for agricultural classification, but zero-shot performance on disease, pest, damage, quality, and species identification remains poor, and it is unclear whether this reflects weak visual features or a failure to connect them to domain knowledge. We build a benchmark of 116 datasets, 834 classes, and 8,324 images spanning these tasks to isolate where the gap arises. Linear probing shows VLM vision encoders already encode agricultural features nearly as separable as a self-supervised DINOv3 baseline, ruling out weak visual representations as the primary bottleneck. Conditioning each model on an oracle reference description (an upper bound on its parametric knowledge) closes most of the gap left by an unaided lower bound, showing VLMs already know more about agriculture than they show. To close this gap without an oracle description at inference time, we structure test-time reasoning around a fixed, per-task diagnostic rubric: the model generates $K$ candidate responses and a Probabilistic Pivot Tournament (PPT) verifier, scored pairwise against the rubric, selects the best one. This nearly doubles judged F1 over the lower bound and matches or exceeds the upper bound on several tasks, notably pushing Gemma 4 E4B-it's disease F1 to 0.71, above its own upper bound of 0.60. However, the verifier's letter-scale confidence score has the opposite of its intended effect: filtering to its most confident predictions does not improve accuracy and correlates negatively with correctness across every model and pool size tested, so the score cannot serve as a measure of predictive uncertainty, and most of the observed gain likely comes from rubric-grounded generation rather than pairwise verification.
CommentsSubmitted to the AI for Science Workshop (NeurIPS Workshops 2026)