发表机构
Duke University; Texas A&M University; Florida State University; UC San Diego; University of Pennsylvania; UNC Chapel Hill(杜克大学; 德克萨斯A&M大学; 佛罗里达州立大学; 加州大学圣迭戈分校; 宾夕法尼亚大学; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GeoSim框架,通过四层几何分析揭示VLM在24个低层视觉任务中表征的组织原则及其跨任务/模型迁移极限。
AI 中文摘要
视觉语言模型(VLM)已成为通用视觉骨干网络的有力候选,代表性架构包括自回归(AR)模型和扩散变换器(DiTs)。然而,如何高效地将它们适配到全合一低层图像恢复任务中仍是一个挑战。关键在于,该领域缺乏对VLM如何组织隐藏层表征以及这些结构上不同的范式是否共享用于像素级感知的公共几何组织的理解。这种共享组织是构建高迁移性、统一恢复VLM和适配器的前提。在本文中,我们系统地研究了涵盖5个类别的24个低层任务中的表征相似性。我们提出了GeoSim,一个统一的四层框架,从全局相似性、局部几何、稀疏特征分解和拓扑验证视角分析任务条件表征。我们的公式适用于分析AR模型中的隐藏状态和DiTs中的特征图,涵盖相同和跨任务/模型设置。我们的结果揭示了低层视觉表征的组织原则,同时暴露了它们在跨任务和跨模型一致性方面的局限性。最终,GeoSim为探测低层视觉中的潜在迁移性以及诊断模型在特定任务或模型场景中的局限性提供了一个可解释性视角。
英文摘要
Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.
CommentsFirst version: 10 pages