arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结的3D CT视觉编码器中的稀疏概念通道

Sparse Concept Channels in Frozen 3D CT Vision Encoders

Farhad Nooralahzadeh, Lea Bogensperger, Christian Bluethgen, Michael Krauthammer

arXiv 2607.20993首次发表:更新:

AI 中文总结

研究3D医学图像解释中视觉语言模型内部单元对临床发现的编码,提出无训练概念通道探测(CCP)方法,通过实验证明该方法能清晰表征冻结医学编码器对发现的表示,在临床疗效和NLG指标上表现优且延迟低。

AI 中文摘要

大型视觉语言模型在3D医学图像解释中日益占主导地位,但我们很少知道哪些内部单元编码临床发现以及该信息在表示中的位置。我们首先通过探测其冻结的视觉嵌入在3D胸部视觉语言模型(Pillar-0)上进行研究。结果表明:每个放射学发现由约10个视觉编码器通道的稀疏集编码,关闭与一个发现相关的通道会使该发现的分数下降,相同的稀疏探测在架构不相关的3D腹部VLM(Merlin)上也能复制。我们的无训练概念通道探测(CCP)方法在临床疗效和NLG指标上优于已发表的CT-CHAT,延迟降低22倍。我们的结果清晰、可重复地表征了冻结的医学编码器如何表示发现,证明了在模型间的直接适用性。

英文摘要

Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know <i>which</i> internal units encode clinical findings or <i>where</i> that information lives in the representation. We first study this on a 3D chest vision-language model (Pillar-0) by probing its frozen vision embeddings. We show that (i) each radiological finding is encoded by a <i>sparse</i> set of ~10 vision-encoder channels that match full-feature classification performance and far exceed a zero-shot text prompting; (ii) turning off the channels tied to one finding, that finding's score collapses while unrelated labels stay stable; and (iii) the same sparse probe <i>replicates</i> on an architecturally unrelated 3D abdominal VLM (Merlin) suggesting a general property of frozen medical encoders. Our training-free concept channel probe (CCP) method, paired with a corpus-derived report template, outperforms published CT-CHAT on clinical efficacy and NLG metrics (F1 0.549 vs. 0.184; BLEU 0.483 vs. 0.373) at 22x lower latency. Our results provide a clear, reproducible characterization of how frozen medical encoders represent findings, demonstrating direct applicability across models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑