多模态大语言模型时代的容积放射学人工智能
Volumetric Radiology AI in the Era of Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
该综述梳理截至2026年7月的200余篇文献,探讨多模态大语言模型时代容积放射学AI的表征、智能体系统及临床评估,提出相关框架以明确原生容积建模的适用场景。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)的进展正将放射学人工智能(AI)从特定任务的图像分析拓展至多模态理解与推理。然而容积放射学存在根本的表征不匹配问题:临床解读常需全容积空间上下文及与采集相关的定量信息,而当前MLLMs通常以选定的二维(2D)图像、压缩的视觉表征或报告衍生文本为条件。因此可靠的容积放射学AI需要保留任务相关三维(3D)信息的表征,以及能在临床工作流程中访问、验证和整合该信息的系统。本综述梳理了截至2026年7月的200余篇文献,从模型层面的容积表征与多模态理解、系统层面的智能体编排,及其与临床应用和评估的关联对文献进行组织。综述涵盖容积基础模型、语言对齐与压缩策略,以及通过规划、工具、记忆和工作流程交互扩展MLLMs的智能体系统;区分选定2D视图或报告介导推理即可满足需求的场景,与需原生容积建模的场景;还引入Claim-Design-Validation框架,评估技术、工作流程和临床主张是否与恰当的设计及验证相匹配。文献显示,原生容积建模与智能体能力取决于目标任务的空间、定量、上下文及工作流程需求;临床可信度需在真实工作流程中具备忠实的容积表征、可追溯的系统行为、与主张对齐的验证,以及明确的人类监督。
英文摘要
Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.
发表机构
- Fudan University(复旦大学)
- Northwestern Polytechnical University(西北工业大学)
- Xi’an Jiaotong University(西安交通大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Southern Medical University(南方医科大学)
- The Chinese University of Hong Kong(香港中文大学)
- University of British Columbia(不列颠哥伦比亚大学)
- Microsoft Research(微软研究院)
- Westlake University(西湖大学)
机构由 AI 辅助整理,请以论文原文为准。