arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14499cs.AI

通过动态多轮交互对视觉语言模型进行情境化评估

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

  • UC San Diego(加州大学圣地亚哥分校)
  • Northeastern University(东北大学)
  • Georgia Institute of Technology(佐治亚理工学院)
  • Johns Hopkins University(约翰·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao

AI总结:

研究多模态大语言模型现实有效性问题,提出CEDI框架,通过三方交互、多轮半结构化对话及多种策略评估,应用于视觉幻觉,发现能揭示更多接近实际情况的幻觉,凸显其对MLLMs能力评估的作用。

AI中文摘要:

多模态大语言模型(MLLMs)在基准测试中取得了显著进展,但其在现实世界中的有效性仍不确定。这种差距源于受控静态环境中的基准测试与现实世界应用的动态、交互和情境化性质之间的根本错位。为弥合这一差距,我们提出了CEDI(通过动态多轮交互对MLLMs进行情境化评估)框架,将评估重新构建为被评估模型、自动考官和评分者之间的三方交互。考官通过基于任务的图形表示进行多轮半结构化对话。通过导航状态空间转换,CEDI部署从澄清请求到对抗性探测等各种策略,以获取性能证据。我们将CEDI应用于视觉幻觉。多个模型、不同设置、数据集和领域的实证结果表明,情境化、交互式评估不仅比传统静态评估揭示出更多幻觉,而且揭示出的幻觉更接近实际用例中出现的幻觉。我们还表明,幻觉往往会通过自我强化的对话历史在长语境中累积,并且模型特别容易受到需要拒绝前提或拒绝的问题的影响。这些发现共同凸显了CEDI是朝着对MLLMs能力进行现实、系统和生态有效评估迈出的一步。

英文摘要:

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs' capabilities. Code is available at github.com/williamium3000/cedi.

↑