AI 中文总结
针对3D医学VLMs的视觉令牌冗余压缩问题,本文提出无训练框架MedARC,整合三类线索估计令牌重要性,通过显著性感知合并策略压缩令牌,在CT-RATE和MR-RATE上实现性能保持的同时降低开销。
AI 中文摘要
将3D医学图像与视觉-语言模型(VLMs)相结合,对计算机辅助诊断具有重要应用前景。然而,体积图像会产生过长的视觉令牌序列,且存在大量空间及切片间冗余。现有令牌压缩方法通常采用均匀缩减或依赖单一重要性信号,存在移除与查询相关的临床区域或结构独特区域的风险。为解决这一局限,本文提出MedARC,一种面向3D医学VLMs视觉令牌的统一无训练自适应冗余压缩框架。MedARC通过整合三类互补线索估计令牌重要性:VLM视觉编码器的自注意力(反映模型内在视觉焦点)、投影视觉令牌与文本嵌入的相似度(识别查询相关区域)、局部视觉基础模型特征与体积级特征中心的偏差(突出结构独特的解剖结构)。所得重要性分布引导显著性感知合并策略,在保留信息令牌的同时合并冗余令牌,而非直接丢弃。在CT-RATE和MR-RATE上的实验表明,MedARC可降低视觉令牌开销与推理时间,同时保持或提升诊断性能;其多线索评分的成本远低于处理更少令牌带来的节省,对于更大规模语言模型预计将产生更显著的效益。
英文摘要
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. Existing token compression methods typically apply uniform reduction or rely on a single importance signal, increasing the risk of removing regions that are clinically relevant to the query or structurally distinctive. To address this limitation, we propose MedARC, a unified, training-free framework for Adaptive Redundancy Compression of visual tokens in 3D medical VLMs. MedARC estimates token importance by integrating three complementary cues: self-attention from the VLM vision encoder, which reflects the model's intrinsic visual focus; similarity between projected visual tokens and text embeddings, which identifies query-relevant regions; and deviations of local visual foundation model features from the volume-level feature center, which highlight structurally distinctive anatomy. The resulting importance distribution guides a saliency-aware merging strategy that preserves informative tokens while consolidating redundant ones rather than simply discarding them. Experiments on CT-RATE and MR-RATE show that MedARC reduces visual-token overhead and inference time while preserving or improving diagnostic performance. Its multi-cue scoring cost is outweighed by the savings from processing fewer tokens, with greater benefits expected for larger language models.
Comments9 pages