arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30928cs.CVcs.AI

UltraG-Bench:用于评估大型视觉语言模型在超声像素级证据定位上的多任务基准

UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao, Hongfei Lin, Feng Xia

首次发表
浏览论文内容

中文总结 AI 辅助

针对超声图像理解中缺乏像素级证据定位的问题,提出多任务基准UltraG-Bench及结合VLM与UltraSAM3的UltraG-Agent,显著提升语义预测与像素级定位能力。

中文摘要 AI 辅助

超声是最广泛使用的医学成像方式之一,近年来大型视觉语言模型(VLMs)在超声图像理解方面展现出日益增强的能力。然而,这些模型无法提供与其语义预测对齐的像素级视觉证据,且它们在超声中的细粒度定位能力在很大程度上仍不清楚。我们引入了UltraG-Bench,一个用于评估超声中像素级证据定位的大规模多任务基准。UltraG-Bench通过标注40个公共超声分割数据集构建,涵盖13个解剖类别,并包含三个递进任务:指令引导分割、证据接地VQA和证据接地报告生成,分别具有331125、666779和138832个标注。对14个最先进模型的综合评估揭示了语义理解与细粒度像素级定位之间的显著差距。我们进一步提出了UltraG-Agent,它结合了VLM的语义推理能力与UltraSAM3的超声特定分割能力。实验表明,UltraG-Agent显著提升了语义预测和像素级视觉定位。我们的数据集和代码可在该https URL获取。

英文摘要

Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.

发表机构

  • Dalian University of Technology(大连理工大学)
  • RMIT University(皇家墨尔本理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑