arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于多模态大语言模型评估的深度交错文本-图像上下文

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo

arXiv 2609.02573首次发表:更新:

发表机构

Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对多模态模型评估中忽略深度交错文本-图像场景的问题,提出基准 TIC-Bench,评估 10 种多模态大语言模型的相关能力,发现其与人类专家存在显著性能差距。

AI 中文摘要

当前多模态模型的评估与训练主要聚焦于多图像任务,很大程度上忽略了文本-图像交错的场景。在这类多图像任务中,文本通常仅作为任务指令,缺乏与视觉内容的深度语义交互。相比之下,文本-图像协同创作、角色跟踪、空间重建等实际应用需要文本与图像之间的持续交互,因此模型必须具备对这些交错上下文的深度理解能力。为弥合这一差距,本文提出了一种新的基准 TIC-Bench(深度交错文本-图像上下文),用于评估模型整合文本-图像线索并在深度交错上下文中恢复真实事实的能力。该基准涵盖逻辑关联、时间关联、空间关联三个核心领域,进一步细分为八种具体类型,共包含 2280 个问题。本文评估了 10 个最先进的多模态大语言模型(MLLMs),观察到它们与人类专家相比存在显著的性能差距,且始终难以整合分布在交错视觉和文本输入中的证据。最终,该基准为评估和提升多模态模型在深度交错上下文中有效整合文本与图像信息的能力提供了宝贵的分析工具,TIC-Bench 可通过此公开链接获取。

英文摘要

Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑