P3CA:通过空间探测实现与编码器无关的视觉基础模型嵌入解释
P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing
- School of Computing, Queen’s University(女王大学计算学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视觉基础模型嵌入难以解释的问题,提出与编码器无关的P3CA方法,通过空间提示实现局部探测,在多类数据上验证其可揭示局部结构、提升病理判别能力等贡献。
AI中文摘要:
视觉基础模型正越来越多地被用作医学图像计算中的可复用编码器,然而除了下游任务性能或全局降维外,其高维空间嵌入难以被检查。我们提出了位置提示主成分分析(P3CA),这是一种与编码器无关的、针对通道丰富的空间张量的局部探测方法。给定用户选定的空间提示,P3CA会估计该区域内的特征归一化和主导协方差方向,随后将得到的投影应用于整个张量,以可视化局部信息方向的表达位置。这产生了一种区域条件表征透镜,无需修改编码器、重新训练或特定任务标签。我们在EmbedVision(一种基于3D Slicer的交互式工作流)中实现了P3CA,并在自然图像、结直肠癌病理基础模型嵌入以及空间转录组张量上对其进行评估。在这些设置中,提示投影揭示了被全局主成分分析抑制的局部结构,从冻结的三维投影中提升了与提示匹配的病理判别能力,并支持学习到的与测量得到的空间表征之间的比较。
英文摘要:
Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are difficult to inspect beyond downstream task performance or global dimensionality reduction. We propose position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors. Given a user-selected spatial prompt, P3CA estimates the feature normalization and dominant covariance directions within that region, then applies the resulting projection to the full tensor to visualize where locally informative directions are expressed. This produces a region-conditioned representation lens without modifying the encoder, retraining, or requiring task-specific labels. We implement P3CA in EmbedVision, an interactive 3D Slicer-based workflow, and evaluate it across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors. Across these settings, prompted projections reveal local structure suppressed by global PCA, improve prompt-matched pathology discrimination from frozen three-dimensional projections, and support comparison between learned and measured spatial representations.