幻觉的所在:VQ标记视觉语言模型中的跨架构电路
Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究通过激活修补识别出VQ标记VLM中共享的早期层注意力路由电路,提出三门诊断方法,并证明仅$L_0$消融可显著减少开放式生成中的物体幻觉,揭示幻觉源于架构与预训练。
中文摘要 AI 辅助
通过向量量化(VQ)码本对图像进行标记的统一视觉语言模型(VLM)在基于事实的是/否基准测试中经常产生物体幻觉,然而现有的解码时修复方法将其视为一般的校准错误,缺乏架构层面的解释。通过对跨越八个LLM家族的二十五个模型进行激活修补,我们识别出一个在VQ标记VLM中共享的早期层($L_0$)注意力路由电路,并提出一个三门诊断方法,用以区分携带该电路的模型与不携带该电路的模型。该诊断方法识别出十个阳性模型(五个跨三个LLM家族的自然统一VQ-VLM和五个诱导变体),并排除其余十五个模型。一个单变量架构替换(LLaVA-1.6 CLIP+MLP $\ ightarrow$ VQ+Linear)即可安装该电路,而在相同数据上使用匹配计算量的MLP对照则不会,从而将向量量化隔离为病理信号的来源;承载该信号的路由通路是骨干网络已经提供的。与调优的VCD和DoLA基线相比,调优的DoLA在二元校准上胜出,但**只有$L_0$消融能减少开放式生成中的物体幻觉**(CHAIR$_i$相对减少$31\%$,而调优的DoLA和VCD则保持不变或使其恶化)。这些结果将统一VQ VLM中的物体幻觉重新定义为架构和预训练的特性,并产生了一种机制无关解码无法复制的针对性干预。
英文摘要
Unified vision-language models (VLMs) that tokenize images through a vector-quantized (VQ) codebook routinely hallucinate objects on grounded yes/no benchmarks, yet existing decoding-time fixes treat this as generic miscalibration without an architectural account. Using activation patching across twenty-five models spanning eight LLM families, we identify an early-layer ($L_0$) attention routing circuit shared across VQ-tokenized VLMs and propose a three-gate diagnostic that distinguishes the models carrying it from those that do not. The diagnostic isolates ten positive models (five natural unified-VQ VLMs across three LLM families and five induced variants) and rejects the remaining fifteen. A single-variable architectural swap (LLaVA-1.6 CLIP+MLP $\rightarrow$ VQ+Linear) installs the circuit, while a matched-compute MLP control on identical data does not, isolating vector quantization as the source of the pathological signal; the routing pathway that carries it is one that the backbone already provides. Against tuned VCD and DoLA baselines, tuned DoLA wins on binary calibration, but \textbf{only $L_0$ ablation reduces object hallucination in open-ended generation} (CHAIR$_i$ reduces by $31\,\%$ relatively, whereas tuned DoLA and VCD leave it unchanged or worsen it). These results recast object hallucination in unified VQ VLMs as a property of architecture and pretraining, and yield a targeted intervention that mechanism-agnostic decoding cannot replicate.
发表机构
- Adobe
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。