arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你即你所提示的:农业食品视觉语言模型中的提示质量、领域偏移与不确定性

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

Andrea Morales-Garzón, Salvador López-Joya, Miguel López-Pérez, Maria J. Martin-Bautista

arXiv 2608.18116首次发表:更新:

发表机构

University of Granada(格拉纳达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对农业食品领域,评估了零样本提示集成(ZPE)在分布内与分布外场景的表现,提出PID方法提升严重领域偏移下的故障检测能力,验证了领域特定提示池的优势。

AI 中文摘要

视觉语言模型可通过自然语言提示实现零样本分类,但其性能对提示的构建方式十分敏感,尤其在专业领域中。零样本提示集成(ZPE)通过判别性信号对提示进行加权来解决该问题,然而其在领域偏移下的表现尚未得到探究。我们在农业食品领域中,使用CLIP和SigLIP在四个数据集和四个提示池中对ZPE进行评估,涵盖分布内(ID)食品和分布外农业基准。ZPE在分布内条件下提供的益处有限,但在领域偏移下可显著提升性能与校准效果,其中51至52个提示的特定领域池始终优于247至426个提示的通用池。词汇分析显示,ZPE可作为无标签访问的无监督领域对齐检测器。我们进一步引入PID(基于提示的不一致检测),将提示分歧重新用作认知不确定性,在标准置信度度量失效的严重领域偏移下提升故障检测能力。

英文摘要

Vision-language models enable zero-shot classification through natural language prompts, but performance is sensitive to prompt formulation, especially in specialized domains. Zero-shot Prompt Ensembling (ZPE) addresses this by weighting prompts by discriminative signal, yet its behavior under domain shift remains unexplored. We evaluate ZPE in the agrifood domain using CLIP and SigLIP across four datasets and four prompt pools, spanning in-distribution (ID) food and out-of-distribution agricultural benchmarks. ZPE provides limited benefit under ID conditions but substantially improves performance and calibration under domain shift, where domain-specific pools of 51-52 prompts consistently outperform generic pools of 247-426. Lexical analysis shows that ZPE acts as an unsupervised domain-alignment detector without label access. We further introduce PID (Prompt-based Inconsistency Detection), which repurposes prompt disagreement as epistemic uncertainty, improving failure detection under severe domain shift where standard confidence measures collapse.

CommentsAccepted in the journal Procesamiento del Lenguaje Natural (SEPLN2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑