发表机构
University of Melbourne; Johns Hopkins Applied Physics Lab; Johns Hopkins University; Human Language Technology Center of Excellence; Monash University(墨尔本大学; 约翰斯·霍普金斯大学应用物理实验室; 约翰斯·霍普金斯大学; 人类语言技术卓越中心; 莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LAYERSCOPE提出无标签逐层几何度量框架,用于比较视频和多模态模型表示,发现中间层优于最终层,且RankMe对分类聚类最有效。
AI 中文摘要
我们提出了LAYERSCOPE,一个无标签、逐层的框架,旨在刻画模型在视频和多模态设置中学到的表示。使用最终层或中间层的表示来评估下游性能通常需要大量带标签的数据、重复的任务特定评估以及大量的计算。为解决这些限制,LAYERSCOPE使用局部、全局、分布性和基于对应关系的几何度量来比较模型内部及跨模型的逐层表示结构,而无需任务特定的标签。我们在MVEB/MVEB+上评估了七个架构多样的模型,涵盖视频和多模态分类、聚类以及文本到视频检索任务。我们发现中间层表示可以优于最终层和模型默认输出。我们还发现,没有任何单一的几何度量能一致地预测下游性能,但注意到不同模型系列中出现了不同的逐层几何特征。LID显示出与性能的任务相关关系,而RankMe为分类和聚类提供了最强的度量,但并非通用的层选择器。我们还发现,配对感知度量比仅使用分布性距离能更好地解释检索性能。因此,LAYERSCOPE提供了一个跨模型和跨层比较表示的框架,使得在视频和多模态设置中能够进行更系统的评估。
英文摘要
We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated task-specific evaluations, and substantial computation. To address these limitations, LAYERSCOPE uses local, global, distributional, and correspondence-based geometric metrics to compare layerwise representation structure within and across models without requiring task-specific labels. We evaluate seven architecturally diverse models across video and multimodal classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs. We also find that no single geometric metric consistently predicts downstream performance, but note that distinct layerwise geometric signatures emerge across model families. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that pairing-aware metrics explain retrieval better than distributional distances alone. LAYERSCOPE therefore offers a framework for comparing representations across models and layers, enabling a more systematic evaluation in video and multimodal settings.
CommentsPreprint, minor corrections