arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27511cs.CVcs.AI

NV-Reason-CT:用于CT分析的3D视觉语言模型

NV-Reason-CT: 3D Visual Language Model for CT Analysis

发表机构英伟达 · 美国国立卫生研究院 · 牛津大学
另 4 家 · 查看机构详情
  • NVIDIA(英伟达)
  • NIH(美国国立卫生研究院)
  • University of Oxford(牛津大学)
  • CHOP/UPenn(费城儿童医院/宾夕法尼亚大学)
  • Basaksehir Cam and Sakura City Hospital(巴萨克谢希尔卡姆和樱市医院)
  • Forithmus
  • University of Zurich(苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, S… 展开作者

Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Zongwei Zhou, Wenxuan Li, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu

首次发表
浏览论文内容

中文总结 AI 辅助

提出NV-Reason-CT,一种结合原生3D视觉编码与放射科医生引导推理的生成式视觉语言模型,用于胸腹CT的异常分类、报告生成和交互式推理,在CT-RATE上取得高F1和AUROC,并显著减少报告时间。

中文摘要 AI 辅助

我们提出了NV-Reason-CT,一个用于胸部和腹部CT的生成式视觉-语言模型,该模型将原生3D视觉编码与放射科医生引导的推理相结合。该模型将原生3D视觉变换器与语言模型耦合,将所有视觉标记及其显式3D坐标传递到语言解码中,而无需进一步的空间标记合并。这保留了视觉编码器内的体积空间信息,并在与文本联合处理期间通过语言模型的位置编码保留该信息。我们在大约550,000个多模态指令示例的精选语料库上训练,这些示例来自70,111个独特的CT图像输入,结合了标准化报告、异常聚焦和解剖学特定问题、多轮交互以及来自记录和转录的专家CT解释的放射科医生撰写的推理。专家注释提供直接监督,并指导额外的报告基础的合成推理。端到端监督微调(SFT)之后是组相对策略优化(GRPO),在胸部和腹部异常集上具有可验证的奖励。该模型支持异常分类、报告生成以及具有可审查观察结果、鉴别诊断和不确定性的交互式推理。评估涵盖公共CT基准和一个保留的NIH队列。在CT-RATE上,NV-Reason-CT在没有任务特定分类头的情况下实现了宏F1为0.614和宏AUROC为0.871;生成的报告实现了报告派生的宏F1为0.592。在与专家放射科医生的初步研究中,AI辅助审查获得了有利的置信度评级,并与平均报告解释和报告时间减少50%相关。我们发布模型和训练代码,以支持体积医学成像的可解释AI的可重复研究。

英文摘要

We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.

↑