(如何)多模态大语言模型像人类一样报告双稳态图像?
(How) Do MLLMs Report Bistable Images Like Humans?
浏览论文内容
中文总结 AI 辅助
本研究探讨多模态大语言模型是否像人类一样报告双稳态图像,发现其报告可被视觉和语言线索调节且具排他性,机制涉及竞争性图像表征和不同调节路径。
中文摘要 AI 辅助
双稳态图像(如鸭兔图)是经典刺激,其中一幅图像支持多种互不兼容的解释,人类通常一次只报告一种。我们探究多模态大语言模型(MLLMs)是否表现出类似的报告行为,以及何种内部计算支持这种行为。利用LLaVA系列模型,我们研究两个可处理的维度:可调节性,即报告是否可被自下而上的视觉线索和自上而下的语言先验所偏向;以及排他性,即响应是否承诺单一解释。我们在经典的鸭兔图和合成的视觉错位图(Visual Anagrams)上测试这两点,以减轻记忆混淆的影响。在行为上,视觉和语言操纵均以与人类一致的方式系统性改变报告,而响应仍主要保持排他性。在机制上,这些效应源于竞争性的图像令牌表征、自下而上和自上而下调节的不同路径,以及排他性报告与物体计数编码之间的联系。代码和数据可在该https URL获取。
英文摘要
Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both on the canonical duck-rabbit and on synthetic Visual Anagrams to mitigate memorization confounds. Behaviorally, both visual and linguistic manipulations systematically shift reports in human-consistent ways, while responses remain predominantly exclusive. Mechanistically, these effects arise from competing image-token representations, distinct pathways for bottom-up and top-down modulation, and a link between exclusive reporting and object-count encoding. Code and data are available at https://github.com/rtakatsky/mllm-bistable-images.
发表机构
- University of Sussex(萨塞克斯大学)
- The University of Tokyo(东京大学)
- RIKEN(理化学研究所)
- Canadian Institute for Advanced Research(加拿大高等研究院)
- Tohoku University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。