通过带可解释生成控制的吉布斯采样探测多模态大语言模型(MLLMs)的感知先验
Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls
浏览论文内容
中文总结 AI 辅助
本文提出带可解释生成控制的吉布斯采样方法,直接从多模态大语言模型的感知先验分布采样,探测出规范偏差及直接提示不可见的新颖先验,为研究模型感知先验提供了新途径。
中文摘要 AI 辅助
模型在任务上的行为由其接收的输入与自身携带的先验(即其隐含期望的刺激分布)共同决定。传统可解释性研究通过固定输入并检查模型响应来研究模型:要么从机制层面探测内部结构如何表征输入,要么从行为层面测量输入变化如何导致输出变化。但这两种方式均无法重构先验分布本身——内部结构仅展示模型能表征的内容,而非其期望的内容;且任何固定刺激集都会使大部分可能的输入空间未被观测到,现实场景中视觉语言模型(VLMs)接触的图像等输入空间维度极高、多样性极强,这些先验仍是影响模型现实行为却鲜为人知的组成部分。本文提出一种方法,通过引导生成模型沿可控轴生成刺激,并以研究目标模型为评判者在该空间上运行吉布斯采样,直接从模型的感知先验分布中采样。我们将该方法应用于多种类别和目标变量(如人脸的可信度、艺术图像的廉价度),既恢复了规范偏差,也发现了直接提示不可见的新颖先验,值得进一步研究其下游效应。
英文摘要
A model's behavior on a task is jointly determined by the input it receives and the prior it brings in, i.e. the distribution over stimuli it implicitly expects. Interpretability research has traditionally studied models by holding inputs fixed and examining model responses either mechanistically, probing how internal structure represents inputs, or behaviorally, measuring how variation in inputs leads to variation in outputs. Neither reconstructs the prior distribution itself, since internal structure shows what a model can represent, not what it expects, and any fixed stimulus set leaves most of the possible input space unseen. In particular, such an input space in real-world settings, such as images seen by VLMs, is extremely high-dimensional and diverse. These priors thus remain a poorly understood component of models that nonetheless influence real-world behavior. We propose a method to sample from models' perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge. We apply this to a variety of categories and target variables (such as trustworthiness in faces and cheapness in art images) and recover both canonical biases and surprising novel priors invisible to direct prompting, warranting further investigation of their downstream effects.
发表机构
- MIT(麻省理工学院)
- Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。