视觉-语言模型是脆弱的多语言关联器
Vision-Language Models are Fragile Multilingual Associators
浏览论文内容
中文总结 AI 辅助
该研究针对视觉-语言模型的多语言关联稳定性问题,构建M²BIND基准,发现其跨语系/文字时会出现绑定崩溃,关系密切语言则表现较好,表明多语言部署的VLMs关联质量存疑。
中文摘要 AI 辅助
视觉-语言模型必须将视觉实体与文本属性相关联。当输入的语言发生变化时,这些关联或概念绑定是否保持稳定尚未被探索。我们引入M²BIND,这是一个在多种语言中改变上下文和查询语言的基准。我们通过任务性能指标从外在角度评估绑定,并通过因果干预从内在角度评估绑定。我们发现绑定不具有语言不变性:跨语系和跨文字设置会触发显著的绑定崩溃,模型的内部绑定计算会转移到更靠后的层并失去因果强度。关系密切的语言能相对更好地保留关联。从更广泛的意义上说,我们的发现表明,在多语言环境中部署的视觉-语言模型不能被假设保持在单语言评估中观察到的相同关联质量。
英文摘要
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language-invariant: cross-family and cross-script settings trigger significant binding collapse, with the model's internal binding computation shifting to later layers and losing causal strength. Closely related languages preserve associations comparatively better. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation.
发表机构
- Manipal University Jaipur(斋浦尔马尼帕尔大学)
- University of North Carolina Charlotte(北卡罗来纳大学夏洛特分校)
- University of Salford(索尔福德大学)
- University of Manchester(曼彻斯特大学)
- Indian Statistical Institute Kolkata(印度统计研究所加尔各答分所)
机构由 AI 辅助整理,请以论文原文为准。