AI 中文总结
研究多模态大语言模型生成图像字幕时的系统性对齐错误检测问题,提出利用现成基础模型的结构化双阶段设置的Symbal方法,引入SymbalBench基准,Symbal在基准上表现出色,可辅助审核MLLM生成的字幕。
AI 中文摘要
多模态大语言模型(MLLMs)生成图像字幕时常出错,导致图像与文本对不对齐。本文聚焦一类系统性对齐错误,即MLLM生成字幕中的反复错误与配对图像中特定视觉特征相关。给定含MLLM生成字幕的视觉语言数据集,目标是检测此类错误。提出Symbal,利用现成基础模型的结构化双阶段设置识别错误并用自然语言总结结果。还引入SymbalBench基准,由两个领域的170万图像文本对组成420个视觉语言数据集。Symbal在基准上表现出色,能准确检测字幕中的系统性对齐错误,可辅助审核MLLM生成的字幕。
英文摘要
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present Symbal, which utilizes a structured, dual-stage setup with off-the-shelf foundation models to identify systematic misalignments and summarize results in natural language. As our second key contribution, we introduce SymbalBench, a benchmark designed to evaluate automated methods on our proposed task. SymbalBench consists of 1.7 million image-text pairs from two domains (natural and medical images), organized into 420 vision-language datasets with annotated systematic misalignments. Symbal exhibits strong performance on this benchmark, correctly identifying systematic misalignments in 63.8% of datasets, a nearly 4x improvement over the closest baseline. We supplement our evaluations on SymbalBench with real-world evaluations, showing that (1) Symbal can accurately surface systematic misalignments in captions generated by four MLLMs and (2) Symbal is a powerful tool for auditing off-the-shelf image-caption datasets. Ultimately, our novel task, method, and benchmark can aid users with auditing MLLM-generated captions and identifying critical errors, without requiring access to the underlying MLLM. Code is available at https://github.com/Stanford-AIMI/Symbal.
CommentsICML 2026