arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

廉价探针预测3D-CT视觉语言模型中的昂贵训练

When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation

Renjie Liang, Zijian Xu, Jinqian Pan, Chengkun Sun, Zhengkang Fan, Shawn Li, You Qin, Mei Liu, Jie Xu

arXiv 2607.22771首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究为3D-CT视觉语言模型选编码器及压缩方案时,候选组合多难比较。提出用廉价探针替代,构建含验证门限的探测基准,比较读出头并配对探针与下游任务,早期结果显示廉价探针排序与昂贵微调一致,有望快速筛选方案。

AI 中文摘要

为3D-CT视觉语言模型选择冻结图像编码器及其之上的令牌压缩方案,需要在众多候选方案中进行搜索。有多种编码器、压缩令牌的方法和令牌预算,组合增长迅速。通常比较它们意味着在每个组合上微调大语言模型,这需要大量计算。我们探讨廉价探针能否替代这种比较。构建了基于(编码器×压缩)单元的图像基础探测基准,有临床属性家族和两个验证门限。在此基准上比较了一系列读出头,并在初步研究中将每个探针与其匹配的下游任务配对。早期信号令人鼓舞:廉价探针对候选方案的排序与昂贵的微调高度一致,目前在测量的单元上约为r≈0.95。若成立,可在几分钟内用冻结令牌探针筛选编码器和压缩选择,仅对最终方案进行全面训练。

英文摘要

Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap probe on the encoder's representation promises a way out, but whether it forecasts the expensive outcome has never been tested. We test this with CheapCT on report generation and on MeasureVQA, a new VQA dataset we build. MeasureVQA scores the outcome one capability at a time, its answers measured from segmentation masks and Hounsfield units. Report generation scores the whole report at once and reflects mostly disease. CheapCT forecasts expensive training across every capability. The forecast survives changing the probe readout and the language-model backbone. The rank agreement between CheapCT and fine-tuning stays high throughout, from rho = 0.90 to 0.97. Used to choose an encoder, CheapCT picks one nearly as good as the best while fine-tuning a single candidate, at orders of magnitude less compute. We release the code and MeasureVQA at https://github.com/renjie-liang/CheapCT.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑