发表机构
Shenzhen University of Advanced Technology(深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全切片图像给视觉语言模型带来的计算瓶颈,提出解耦路由框架PathSelect,将令牌剪枝重构成序列选择过程,训练时引入方差保持噪声门等,推理时模块分离,有效减少令牌数,提高准确率并优于基于采样的方法。
AI 中文摘要
由于序列长度极长,千兆像素全切片图像(WSIs)给视觉语言模型(VLMs)带来了基本的计算瓶颈。现有方法主要依赖空间采样或无训练剪枝,有稀释微弱但信息丰富信号的风险。我们将WSI令牌剪枝重新表述为一个序列选择过程,提出一个解耦路由框架并集成到预训练的SlideChat基础模型中。引入PathSelect,它采用方差保持噪声门和对角注意力去噪器。在推理时,PathSelect模块完全分离。我们的框架在SlideBench(TCGA)上总体准确率达74.00%,空间令牌减少约36.6倍,优于基于采样的方法。
英文摘要
Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominantly rely on spatial sampling or training-free pruning, which risk diluting weak but informative signals, leading to the loss of critical diagnostic evidence due to the spatially diffuse nature of pathological cues. We reformulate WSI token pruning as a sequential selection process, enabling the model to autonomously learn an optimal routing strategy rather than relying on static heuristics. We herein propose a decoupled routing framework integrated as an active plugin into the fully pre-trained SlideChat base model, leaving both the slide encoder and large language model frozen. To provide continuous gradients for the non-differentiable pruning operation during training, we introduce PathSelect. PathSelect employs a variance-preserving noise gate to modulate each patch's information flow via a differentiable Soft Top-K operator, paired with a diagonal-attention Denoiser that recovers the perturbed representations without semantic leakage. At inference, the PathSelect module is entirely detached. Relying solely on the trained Scorer, a deterministic Hard Top-K operator executes adaptive, data-dependent trajectory termination, significantly accelerating downstream generative processing with exceptionally low sequential token selection latency. Driven by an empirical average of only 44.86 tokens under a maximum constraint of K = 128, our framework achieves 74.00% overall accuracy on SlideBench (TCGA), representing an approximate 36.6x spatial token reduction relative to the uncompressed baseline average while consistently outperforming sampling-based counterparts.