FZ-VLM:用于肺结节特征表征与临床决策的两阶段Florence-Zephyr视觉语言模型框架
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making
另 1 家 · 查看机构详情
- College of Engineering, University of Guelph(圭尔夫大学工程学院)
- Holland Bloorview Kids Rehabilitation Hospital(霍兰德布鲁尔维尤儿童康复医院)
- Guelph General Hospital(圭尔夫综合医院)
- Ontario Veterinary College, University of Guelph(圭尔夫大学安大略兽医学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出首个两阶段视觉语言模型框架FZ-VLM,通过微调Florence-2与Zephyr-7B模型,实现肺CT中肺结节的结构化特征表征与临床决策,性能优于GPT-4基线及人类基线。
中文摘要 AI 辅助
肺癌仍是全球癌症相关死亡的主要原因之一,计算机断层扫描(CT)是筛查和随访评估的主要成像工具。在检测到肺结节后,放射科医生会手动评估解剖位置、直径、边缘特征和衰减类型,以支持风险评估和临床决策制定。然而,这种检测后工作流程耗时,且会受到观察者间差异的影响。现有的人工智能方法通常专注于孤立任务,限制了其作为统一、临床基础的解释框架的应用。本研究提出FZ-VLM,这是一种用于肺CT中统一结构化肺结节特征表征的两阶段Florence-Zephyr视觉语言模型框架。该框架使用微调后的Florence-2模型从专家标注的2D轴位CT切片中提取放射学属性,而Zephyr-7B模型利用这些属性生成结节描述、随访建议和纵向分析。结果显示,第一阶段模型在解剖位置上的准确率为77.18%,边缘特征准确率为67.96%,衰减类型准确率为79.13%,直径估计的平均绝对误差为2.58 mm,优于所评估的基于GPT-4的基线和人类基线。放射科专家对第二阶段的评估显示,准确率为93.9%,完整性评分为98.6%,临床相关性为76.1%,总分为89.5%。安全性分析表明,大多数输出在临床上是安全的,尽管一些随访建议仍需专家审核。据我们所知,本研究提出了首个用于结构化结节特征表征与临床决策的两阶段视觉语言模型框架。
英文摘要
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.