arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37659cs.CVcs.LG

上下文图像值得10个维度吗?

Are In-Context Images Worth 10 Dimensions?

Adhemar de Senneville, Xavier Bou, Jérémy Anger, Rafael Grompone, Gabriele Facciolo

首次发表
浏览论文内容

中文总结 AI 辅助

本文在大型视觉语言模型中发现共享判别几何(SDG),证明早期层通过类似主成分分析的降维机制压缩上下文图像为线性可分离表示,并分析性地展示了线性自注意力实现该过程的原理。

中文摘要 AI 辅助

关于大型语言模型的上下文学习能力已有大量研究,特别是针对归纳电路。对于少样本分类任务,归纳电路利用每个标记样本在上下文中的线性表示来对未标记查询进行分类。然而,很少有研究关注这些线性表示最初是如何构建的。利用视觉模态相比文本的表达性,我们在大型视觉语言模型(LVLMs)中发现了一个共享判别几何(SDG)。这是一个低维空间,跨所有图像分类任务共享,其中上下文图像被压缩为线性可分离的表示,随后用于执行分类。我们观察到这是模型在早期层对视觉表示进行降维的结果。为了解释这一现象:(1)我们分析性地证明线性自注意力可以通过将上下文数据投影到其主成分上来执行降维,每一层实现朝向该目标的一个梯度下降步骤。(2)我们提供证据表明,经过训练的LVLMs通过类似机制在早期层降低视觉表示的维度。

英文摘要

There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.

发表机构

  • Université Paris-Saclay(巴黎-萨克雷大学)
  • CNRS(法国国家科学研究中心)
  • ENS Paris-Saclay(巴黎-萨克雷高等师范学校)
  • Centre Borelli(博雷利中心)
  • Pôle recherche de l’AMIAD(AMIAD研究部)
  • Institut Universitaire de France(法国大学研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑