任务诱导的视觉Transformer特征空间黎曼度量
Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces
浏览论文内容
中文总结 AI 辅助
本文提出任务诱导的黎曼度量用于ViT特征空间,通过谱拉回网络学习低秩近似,并利用诊断指标预测可行性,实现高效token剪枝,显著提升深度估计性能。
中文摘要 AI 辅助
处理视觉Transformer(ViT)特征空间的方法通常依赖于欧氏距离或余弦相似度。这假设每个方向具有同等重要性,但没有理由相信真实的任务几何具有这一性质。特征空间的任务敏感几何由拉回度量$g(F) = J(F)^\ op J(F)$给出,其中$J$是解码器输出(馈送到任务特定距离)相对于特征的雅可比矩阵。在现代规模下存储完整的$g$是不可行的,对于深度图等密集输出,甚至形成$J$也是不切实际的。我们表明,能否学习该度量的低秩近似取决于模型-解码器对,并利用一个无需矩阵的诊断指标$κ_{cap}(r)$(可通过少量雅可比-向量积计算)来刻画这一性质。对于可处理的配对,我们开发了谱拉回网络(SPN),通过随机幂迭代学习度量的低秩版本,并将其蒸馏为一个31万参数的注意力头,直接从特征预测token重要性。当雅可比谱过于分散而无法进行低秩近似时,将解码器的输入特征通过VAE瓶颈可以恢复可处理性。在DPT、DINOv2、CLIP和VGGT骨干网络上,$κ_{cap}(r)$预测了哪些学习度量架构是可行的。注意力头在DINOv2 CLS上达到Spearman $ρ= 0.998$,我们的几何token剪枝在剪枝率0.5下将基于ToMe的token选择的额外深度误差减少了25%,且无需微调ViT。项目页面:https://cyberiada.github.io/TaskInducedViTs/
英文摘要
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason to believe the true task geometry has this property. The task-sensitive geometry of the feature space is given by the pullback metric $g(F) = J(F)^\top J(F)$, where $J$ is the Jacobian of the decoder's output fed to a task-specific distance, with respect to the features. Storing the full $g$ is infeasible at modern scales, and for dense outputs such as depth maps even forming $J$ is impractical. We show that whether a low-rank approximation of this metric can be learned depends on the model-decoder pair, and we characterize this with a matrix-free diagnostic $κ_{cap}(r)$ computable with a low number of Jacobian-vector products. For tractable pairs, we develop the Spectral Pullback Network (SPN), which learns a low-rank version of the metric from randomized power iteration, and we distill it into a $310$K-parameter importance head that predicts token importance directly from the features. When the Jacobian spectrum is too spread out for a low-rank approximation, passing the decoder's input features through a VAE bottleneck can restore tractability. Across DPT, DINOv2, CLIP, and VGGT backbones, $κ_{cap}(r)$ predicts which learned-metric architectures are viable. The importance head reaches Spearman $ρ= 0.998$ on DINOv2 CLS, and our geometric token pruning reduces the additional depth error of ToMe-based token selection by $25\%$ on DPT depth at prune ratio $0.5$, without fine-tuning the ViT. Project page: https://cyberiada.github.io/TaskInducedViTs/
发表机构
- Koç University(科奇大学)
- Imperial College London(伦敦帝国理工学院)
- Hacettepe University(哈斯特帕大学)
机构由 AI 辅助整理,请以论文原文为准。