提示优化器是盲目运行的吗?用于自动提示优化的跨模态视觉反馈
Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
AI总结:
研究多模态任务中自动提示优化的盲目性问题,提出跨模态视觉反馈方法,含故障条件视觉诊断和错误感知聚合阶段,在多个VQA数据集和目标VLM上效果显著,能提升分数且优化器可跨模型转移。
AI中文摘要:
自动提示优化(APO)已被广泛用于在不更新权重的情况下使视觉语言模型(VLM)适应下游任务,并取得了不错的成果。然而,在多模态任务中,APO的有效性从根本上受到盲目反馈通道的限制:优化器读取问题、预测和正确答案,但从不读取模型失败时的输入图像,因此无法诊断基于视觉的错误。作为一种补救措施,我们引入了跨模态视觉反馈(CMVF)。CMVF包括一个故障条件视觉诊断阶段,其中一个更强的优化器VLM在不访问预测或标签的情况下检查每个失败的图像,以及一个错误感知聚合阶段,该阶段将这些观察结果压缩成可重复使用的任务级视觉盲点模式,以驱动提示重写。至关重要的是,图像仅在优化期间被使用;部署的工件是一个普通的文本提示,其推理成本与任何纯文本基线相同。在12个VQA数据集和4个目标VLM上的大量结果表明,CMVF始终排名第一,在每个目标上比最强的基线平均提高2.4分,在个别基准上提高高达6.5分。此外,优化器会自组织成专家式的视觉检查清单,无需重新优化即可跨模型转移。
英文摘要:
Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However, on multimodal tasks, the effectiveness of APO is fundamentally bottlenecked by a blind feedback channel: the optimizer reads the question, the prediction, and the gold answer, but never the input image on which the model failed, and therefore cannot diagnose visually grounded errors. As a remedy, we introduce Cross-Modal Visual Feedback (CMVF). CMVF incorporates (1) a failure-conditioned visual diagnosis stage, in which a stronger optimizer VLM inspects each failed image without access to predictions or labels, and (2) an error-aware aggregation stage that compresses these observations into reusable, task-level visual blind-spot patterns that drive the prompt rewrite. Crucially, the image is consumed only during optimization; the deployed artifact is an ordinary text prompt that runs at the same inference cost as any text-only baseline. Extensive results across 12 VQA datasets and 4 target VLMs demonstrate that CMVF consistently ranks first, improving over the strongest baseline on every target by 2.4 points on average, with gains of up to 6.5 points on individual benchmarks. Moreover, the optimizer self-organizes into expert-style visual checklists that transfer across models without re-optimization.