基于视觉语言模型(VLM)与大语言模型(LLM)驱动的多智能体正电子发射断层扫描(PET)图像去噪系统
VLM- and LLM-Driven Multi-Agent System for PET Image Denoising
- University of Florida(佛罗里达大学)
- University of Bern(伯尔尼大学)
- University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对PET图像分辨率低、信噪比差及深度学习去噪部署需多模型与专家干预的问题,本文提出VLM与LLM驱动的多智能体闭环PET去噪框架,自主选最优模型参数,在低剂量数据上优于UNet等基线方法。
AI中文摘要:
正电子发射断层扫描(PET)成像存在空间分辨率有限、信噪比低的问题,会损害定量准确性与病灶可检测性。基于深度学习的去噪方法在提升PET图像质量方面展现出强大潜力,但在实际场景中的部署仍具挑战,通常需要多个专用模型与专家干预,例如识别运动诱发的配准伪影、估计噪声水平以选择合适的去噪器,以及去噪后进行病灶聚焦的定量评估。近期,用于图像质量理解的视觉语言模型(VLM)和用于上下文推理的大语言模型(LLM)的进展,为自动化、决策驱动的工作流提供了新机遇。受PET图像质量增强的专家工作流启发,本文提出一种VLM与LLM驱动的多智能体PET去噪框架,可动态评估图像质量与病灶状态,自主选择最优去噪模型与参数,并具备带回滚机制的闭环反馈。实验在Siemens Biograph Vision Quadra PET/CT数据上开展,采用1/20和1/50低剂量设置。各模块评估验证了智能体组件的可靠性,而完整框架在两种剂量水平下均比UNet、GAN和DDPM基线取得更高的峰值信噪比(PSNR)与结构相似性指数(SSIM)。这些初步结果表明,采用闭环多智能体框架使PET去噪策略适配不同图像条件具有可行性。
英文摘要:
Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to-noise ratio, which can compromise quantitative accuracy and lesion detectability. Deep learning-based denoising methods have demonstrated strong potential for improving PET image quality. However, their practical deployment in real-world settings remains challenging, often requiring multiple specialized models and expert interventions, such as identifying motion-induced misregistration artifacts, estimating noise levels to select an appropriate denoiser, and performing lesion-focused quantitative assessment after denoising. Recent advances in vision-language models (VLMs) for image quality understanding and large language models (LLMs) for contextual reasoning provide new opportunities for automated, decision-driven workflows. Inspired by expert workflows for PET image quality enhancement, we propose an VLM- and LLM-driven multi-agent PET denoising framework that dynamically assesses image quality and lesion status, autonomously selects optimal denoising models and parameters, and enables closed-loop feedback with rollback mechanisms. Experiments were conducted on Siemens Biograph Vision Quadra PET/CT data with 1/20 and 1/50 low-dose settings. Individual module evaluations demonstrated the reliability of the agentic components, while the complete framework achieved higher PSNR and SSIM than UNet, GAN, and DDPM baselines at both dose levels. These preliminary results demonstrate the feasibility of using a closed-loop multi-agent framework to adapt PET denoising strategies to different image conditions.