arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26513cs.CV

多模态Transformer中的虚拟编码器

Virtual Encoders in Multimodal Transformers

  • The University of Osaka(大阪大学)

机构由 AI 辅助整理,请以论文原文为准。

Katsuya Ogata, Yuta Nakashima

AI总结:

本文提出虚拟编码器概念,发现多模态Transformer可在内部层自行构建感知表示,无需专用编码器,为理解多模态处理机制提供新视角。

AI中文摘要:

多模态语言模型传统上依赖专用的感知编码器来构建任务可用的表示。近期出现了更集成的架构,这些架构将共享的Transformer直接暴露给轻量投影的补丁、音频帧或离散视觉令牌。当这些表示未被提供时,编码发生在何处?我们发现Transformer可以将这种缺失的计算内化,在其早期到中间层内构建任务可用的感知表示,然后再传递给下游语言模型。我们将这种计算结构称为虚拟编码器。通过线性探测、与感知编码器的相似性以及因果分析,我们在接收感知令牌但无连续编码器派生特征的模型中识别出这种结构的特征。这些分析还表明,感知与语言处理之间的边界不必与架构模块重合。相反,类似编码器的计算可以在共享Transformer内部作为一种功能机制出现,为理解多模态模型在何处以及如何处理感知提供了新视角。

英文摘要:

Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.

↑