发表机构
IIIT Delhi; Dr. Ambedkar Institute of Technology; Heritage Institute of Technology; PwC(德里信息技术学院; 阿姆贝德卡尔技术学院; 遗产技术学院; 普华永道)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型中上下文夹带现象,提出需构建分类结构化双模态工具。引入ENTRAP-VL数据集,含1500个项目,跨八类,分文本和视觉夹带流,提供工具、分类法及评估协议,助力社区严格研究该现象。
AI 中文摘要
上下文夹带是指模型受输入中的辅助上下文影响输出,而不论该上下文是否相关、真实或有意义。近期在单模态语言模型中已被识别并给出机制解释。然而,视觉语言模型(VLM)中是否以及如何表现上下文夹带很大程度上未被研究,该领域也缺乏专门工具来探究。研究VLM中的上下文夹带需要构建一个分类结构化的双模态工具。为此引入ENTRAP-VL,它是一个包含1500个项目、跨八个类别的手动策划数据集,按跨越两个轴的分类法组织,分为文本夹带流(八个上下文条件)和视觉夹带流(三个上下文条件)。我们不测量特定模型中的夹带,而是提供工具、分类法及评估协议,以便社区能严格研究该现象,数据集和文档将公开。
英文摘要
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.