arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15603cs.CVcs.AI

用于PSMA PET/CT报告生成、视觉问答和病灶分割的统一视觉语言模型

A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation

Yang Xing, Jiong Wu, Savas Ozdemir, Yang Zhou, Boxiao Yu, Ying Zhang, Zheren Zhu, Chenyu You, Wei Shao, Yang Lu, Kang Wang, Tinsu Pan, Yang Yang, Kuang Gong

首次发表
浏览论文内容

中文总结 AI 辅助

提出统一PSMA PET/CT视觉语言模型,集成报告生成、视觉问答和病灶分割,在多项指标上优于现有方法,实现多任务一体化分析。

中文摘要 AI 辅助

准确的PSMA PET/CT解读对前列腺癌管理至关重要,然而现有的PET/CT AI模型通常处理孤立的任务。我们提出了一种统一的PSMA PET/CT视觉语言模型,用于报告生成、视觉问答和病灶分割。该框架采用LLaVA风格架构,包括PET/CT视觉编码器、MLP-Mixer投影模块、LoRA微调的大语言模型和3D分割分支。训练遵循四阶段策略:视觉编码器预训练、投影层对齐、VLM微调和最终的多任务微调。语言任务使用了5,747个带有配对报告的PSMA PET/CT数据集,而分割任务使用了AutoPET中的PSMA子集。该模型在标准报告生成指标上优于PET2REP和基于CT的基线,在各类VQA问题上提升了性能,并在Dice和病灶级重叠F1上优于SegAnyPET和nnUNet。这些结果支持了在单一多任务模型架构内实现结构化、交互式、可解释的PSMA PET/CT分析(具有体素级基础)的统一框架的可行性。

英文摘要

Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We propose a unified PSMA PET/CT vision-language model for report generation, visual question answering, and lesion segmentation. The framework adopts an LLaVA-style architecture, comprising a PET/CT vision encoder, an MLP-Mixer projection module, a LoRA-tuned large language model, and a 3D segmentation branch. Training followed a four-stage strategy: vision encoder pretraining, projection-layer alignment, VLM fine-tuning, and final multitask tuning. Language tasks used 5,747 PSMA PET/CT datasets with paired reports, while segmentation used the PSMA subset of AutoPET. The model outperformed PET2REP and a CT-based baseline across standard report-generation metrics, improved performance across VQA question types, and achieved higher Dice and lesion-level overlap F1 than SegAnyPET and nnUNet. These results support the feasibility of a unified framework for structured, interactive, interpretable PSMA PET/CT analysis with voxel-level grounding within a single multitask model architecture.

发表机构

  • University of Florida(佛罗里达大学)
  • University of California, San Francisco(加利福尼亚大学旧金山分校)
  • Stony Brook University(纽约州立大学石溪分校)
  • The University of Texas MD Anderson Cancer Center(得克萨斯大学MD安德森癌症中心)

机构由 AI 辅助整理,请以论文原文为准。

↑