arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FPCO-Dialog:用于评估视觉语言模型中纠正与协作能力的多轮错误前提基准

FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models

Jiayuan Ma, Yuqi Lu, Weiyang Guo, Chenrui Wang, Junyi Shu, Xuebo Liu, Min Zhang, Jing Li

arXiv 2609.03331首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出FPCO-Dialog基准,针对多轮对话中持续错误前提场景,评估20个视觉语言模型的纠正与协作能力,揭示不同模型间的显著差异,相关资源已公开。

AI 中文摘要

视觉语言模型(VLMs)越来越多地被部署在多轮对话场景中,用户可能会用错误的假设描述视觉内容。然而,现有评估很少单独考察当相同的视觉基础错误前提在多轮对话中持续存在时模型的响应情况。我们提出FPCO-Dialog,这是一个用于评估VLMs在重复错误前提下纠正与协作行为的基准。FPCO-Dialog包含1080张图像和10800个问题轮次,按视觉复杂度、物体类别和错误前提类别分层,并采用10轮协议:先有一段正确的对话前缀,随后是重复的错误前提指代表达。我们使用与模型无关的协议和CorrTP@K(针对错误前提轮次的纠正率指标,由两个独立检测器评分)评估了20个商业及开源VLMs。FPCO-Dialog揭示了在基准的替换分布下,不同模型在总体纠正倾向、特定模型的轮次动态以及不同错误前提类型间的系统差异方面存在显著且持续的差异。该数据集、评估协议、模型输出、检测器标签和代码均已公开可用。

英文摘要

Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dialogue turns. We introduce FPCO-Dialog, a benchmark for evaluating correction and cooperation behavior in VLMs under repeated false premises. FPCO-Dialog contains 1,080 images and 10,800 question turns, stratified by visual complexity, object category, and false-premise class, and uses a 10-turn protocol in which a correct dialogue prefix is followed by repeated false-premise referring expressions. We evaluate 20 commercial and open-source VLMs with a model-agnostic protocol and CorrTP@K, a correction-rate metric over false-premise turns, scored by two independent detectors. FPCO-Dialog reveals substantial and persistent cross-model differences in aggregate correction tendency, model-specific turn-wise dynamics, and systematic variation across false-premise types under the benchmark's substitution distribution. The dataset, evaluation protocol, model outputs, detector labels, and code are available.

CommentsAccepted at EMNLP2026 Main Conference. For code and data, see https://github.com/lab-klc/FPCO-Dialog

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑