MinCU:图像对中基于最小变化理解的高细粒度基准
MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs
浏览论文内容
中文总结 AI 辅助
提出MinCU基准和SG-ISA方法,通过隐式空间锚点分解思考-定位-描述序列,提升多模态模型在图像对细粒度变化中的联合描述与定位能力,并降低推理开销。
中文摘要 AI 辅助
定位并描述近乎相同图像之间的细粒度差异,是多模态大语言模型(MLLMs)一项关键但尚未充分探索的能力。现有基准大多孤立地评估语义比较或单图像定位,并未联合要求忠实的描述与物理定位。为弥补这一空白,我们引入了MinCU,一个用于基于最小变化理解的基准,其中每个样本由一对仅在对象类别、属性、数量或空间位置上存在单一原子变化的图像组成,并评估模型描述变化、定位变化区域以及识别变化实体的能力。我们进一步提出了语义引导的隐式空间锚点(SG-ISA),一种结构化的自回归方法,将预测分解为“思考-定位-描述”序列。SG-ISA首先预测变化概念的语义线索,然后使用离散空间锚点作为隐式定位支架,最后生成变化描述及定位框。实验表明,即使是最强的闭源MLLMs和近期R1风格的推理模型在MinCU上也表现挣扎,大多数模型无法联合生成准确的描述和定位框。与之前的思维链方法相比,使用SG-ISA进行微调在定位准确性和描述质量上带来了显著的联合改进,同时将推理令牌开销减少了约26%。这些结果表明,隐式中间空间接口可能比仅依赖模型规模在基于定位的双图像理解中更为有效。
英文摘要
Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and physical localization. To bridge this gap, we introduce MinCU, a benchmark for grounded minimal-change understanding, where each sample consists of an image pair differing by a single atomic variation in object category, attribute, count, or spatial position, and models are evaluated on their ability to describe the change, localize the changed regions, and identify the changed entity. We further propose Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence. SG-ISA first predicts a semantic cue for the changed concept, then uses discrete spatial anchors as an implicit localization scaffold, and finally generates the change description together with the grounding box. Experiments reveal that even the strongest closed-source MLLMs and recent R1-style reasoning models struggle on MinCU, with most failing to jointly produce accurate descriptions and grounding boxes. Compared to the previous chain-of-thought method, fine-tuning with SG-ISA yields substantial joint improvements in grounding accuracy and description quality while reducing reasoning-token overhead by approximately 26%. These results suggest that an implicit intermediate spatial interface can be more effective than relying solely on model scale for grounded dual-image understanding.
发表机构
- Dalian University of Technology(大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。