arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27763cs.CVcs.IRcs.LG

DS@GT ARC参与2026年ImageCLEFmedical任务:医学图像分析中用于概念检测的架构多样性与用于字幕预测的基础模型缩放

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

Bowen Wang, Youwen Zhang, Ritesh Mehta

首次发表
浏览论文内容

中文总结 AI 辅助

该研究介绍DS@GT团队参与2026年ImageCLEFmedical医学图像分析任务,提交不同模型完成概念检测与字幕预测,其中集成模型在概念检测赛道获第一,各模型覆盖不同规模与成本。

中文摘要 AI 辅助

本文介绍了DS@GT团队对2026年ImageCLEFmedical字幕任务的提交情况,该任务是基于ROCOv2数据集的长期基准测试,包含两个赛道:概念检测(任务1),即给放射学图像分配UMLS概念唯一标识符(CUIs);字幕预测(任务2),即生成自然语言字幕。对于任务1,我们的主要提交是ConvNeXt-V2、BiomedCLIP ViT-B/16和DenseNet-169的三路后期融合集成模型,搭配正则化的“诚实阈值调优”程序,旨在避免对稀有概念的验证过拟合;该提交在官方提交中排名第一,主要F1值为0.5790,次要F1值为0.9657。同时,我们提交了一个基于冻结BiomedCLIP嵌入的无训练KNN检索流程,其主要F1值为0.5780,次要F1值为0.9599,在计算成本仅为微调集成模型一小部分的情况下,基本达到了主赛道的性能。对于任务2,我们的提交包括微调后的Gemma-3 27B模型(整体得分0.3571,在官方提交中排名第三)、带有自定义Vizwins合并的全微调BLIP流程(得分0.3564),以及使用PubMed风格提示的零样本MedGemma-4B运行(得分0.3186),覆盖了广泛的模型规模和训练成本。代码:此https URL。

英文摘要

We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.5790$ and a secondary $F_1$ of $0.9657$. In parallel, we submitted a training-free KNN retrieval pipeline over frozen BiomedCLIP embeddings, which reached a primary $F_1$ of $0.5780$ and a secondary $F_1$ of $0.9599$-essentially matching the fine-tuned ensemble on the primary track at a fraction of the cost. For Task 2, our submissions included a fine-tuned Gemma-3 27B model (overall $0.3571$, ranking third in the official submission), a fully fine-tuned BLIP pipeline with custom Vizwins merging ($0.3564$), and a zero-shot MedGemma-4B run with a PubMed-style prompt ($0.3186$), spanning a wide range of model scales and training costs. Code: https://github.com/dsgt-arc/imageclef-caption-2026.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑