发表机构
AI for Good (AIGO); Istituto Italiano di Tecnologia; University of Siena; University of Genoa; University of Verona(人工智能促进美好(AIGO); 意大利技术研究院; 锡耶纳大学; 热那亚大学; 维罗纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出轻量级监督探针OTTER,利用最优传输解码冻结MLLM中修饰词等token的实例定位信息,验证其可扩展至自由生成与跨数据集迁移。
AI 中文摘要
随着多模态大语言模型(Multimodal Large Language Models, MLLM)能够描述日益复杂的视觉场景,token级别的指代定位(grounding)变得至关重要。然而,当MLLM生成“左边的黄色香蕉”时,现有定位方法聚焦于图像中的“什么”(即“香蕉”),却忽略了用于指明具体实例的修饰词token(即“黄色”“左边”)。本研究探究冻结MLLM的表征是否包含可解码的、关于生成token中所指实例的定位信息,该信息可扩展至属性、空间表达、关系/动作术语等修饰词。为解决该问题,我们提出OTTER,一种基于冻结MLLM表征的轻量级监督探针,该探针使用最优传输(Optimal Transport, OT)将生成token与视觉区域对齐,生成紧凑的定位图。实验结果表明:(i)可从修饰词token中解码出实例判别性视觉信息,其中空间术语的证据最为明显;(ii)该信息不仅限于修饰词,上下文对象名词也包含指代信息;(iii)在上下文扰动下,恢复的定位仍具有信息性,所选区域与生成内容保持相关;(iv)所学的基于OT的定位可从受控设置扩展至自由生成及跨数据集迁移。
英文摘要
As Multimodal Large Language Models (MLLMs) can describe increasingly complex visual scenes, token-level grounding becomes crucial. Yet, when an MLLM generates "the yellow banana on the left", established grounding approaches focus on what is in the image ("banana"), overlooking tokens that help describe which instance is meant ("yellow", "left"). In this work, we ask whether frozen MLLM representations contain decodable grounding information about the referred instance across generated tokens, extending to modifiers such as attributes, spatial expressions, and relational/action terms. To address this question, we introduce OTTER, a lightweight supervised probe over frozen MLLM representations that uses Optimal Transport (OT) to align generated tokens with visual regions and produce compact grounding maps. Our results show that (i) instance-discriminative visual information can be decoded from modifier tokens, with the clearest evidence for spatial terms, but (ii) is not confined to them, as contextualized object nouns also carry referential information; (iii) the recovered grounding remains informative under context perturbations, while selected regions remain relevant to generation; and (iv) the learned OT-based grounding extends beyond the controlled setting to free generation and cross-dataset transfer.