基于视觉语言基础模型的文本引导多序列胶质瘤亚区分割细化
Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model
浏览论文内容
中文总结 AI 辅助
该研究开发基于VoxTell的轻量级框架,用3D视觉语言基础模型实现文本引导的胶质瘤亚区分割细化,在BraTS-GLI及跨数据集测试中提升了分割的DSC指标,支持作为临床医生在环工具的进一步评估。
中文摘要 AI 辅助
背景:准确的胶质瘤亚区勾画对放疗计划制定和纵向监测至关重要,但手动轮廓校正耗时。nnU-Net等模型的泛化能力可能存在不足,且缺乏临床医生指导的文本校正。目的:我们研究了将三维(3D)视觉语言基础模型用于文本引导的脑肿瘤分割细化。方法:我们开发了一个轻量级的基于VoxTell的框架。预训练的VoxTell生成初始掩码。源自分割误差的“神谕提示”编码了目标、动作、位置、成像证据、编辑大小和保留约束。冻结的Qwen/VoxTell提示嵌入通过可训练投影注入到其多尺度解码器条件中,其他权重保持冻结。训练、验证和测试使用了901例、100例和250例BraTS-GLI病例。在100例脑膜瘤、转移瘤、儿科肿瘤和UPENN-GBM病例上评估了跨数据集迁移能力。结果:在使用对比后T1加权输入的内部测试集上,正确指令将亚区戴斯相似系数(DSC;增强肿瘤、水肿和坏死/非增强核心)从0.774±0.158提升至0.796±0.137。其性能优于空白提示(0.762±0.155;经Holm校正后p<0.001,dz=0.71)和矛盾提示(0.770±0.163;p<0.001,dz=0.48)。在跨数据集测试中,正确指令将DSC从0.527±0.287提升至0.550±0.278,且优于矛盾指令(0.504±0.275;p<0.001,dz=0.43)。结论:3D视觉语言基础模型可执行指令引导的胶质瘤亚区分割细化。对正确、空白和矛盾提示的敏感性表明这是文本依赖的轮廓编辑,而非非特异性后处理,支持将其作为临床医生在环工具进行进一步评估。
英文摘要
Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such as nnU-Net may generalize imperfectly and lack clinician-directed text correction. Purpose: We investigated adapting a three-dimensional (3D) vision-language foundation model for text-guided brain tumor segmentation refinement. Methods: We developed a lightweight VoxTell-based framework. Pretrained VoxTell generated initial masks. Oracle prompts derived from segmentation errors encoded target, action, location, imaging evidence, edit size, and preservation constraints. Frozen Qwen/VoxTell prompt embeddings were injected through trainable projections into its multiscale decoder conditioning; other weights remained frozen. Training, validation, and testing used 901, 100, and 250 BraTS-GLI cases. Cross-dataset transfer was evaluated on 100 meningioma, metastasis, pediatric tumor, and UPENN-GBM cases. Results: On the internal test set using post-contrast T1-weighted input, correct instructions improved subregion Dice similarity coefficient (DSC; enhancing tumor, edema, and necrotic/non-enhancing core) from $0.774\pm0.158$ to $0.796\pm0.137$. They outperformed blank prompts ($0.762\pm0.155$; Holm-adjusted $p<0.001$, $d_z=0.71$) and contradictory prompts ($0.770\pm0.163$; $p<0.001$, $d_z=0.48$). In cross-dataset testing, correct instructions improved DSC from $0.527\pm0.287$ to $0.550\pm0.278$ and outperformed contradictory instructions ($0.504\pm0.275$; $p<0.001$, $d_z=0.43$). Conclusion: A 3D vision-language foundation model can perform instruction-guided refinement of glioma subregion segmentations. Sensitivity to correct, blank, and contradictory prompts suggests text-dependent contour editing rather than nonspecific post-processing, supporting further evaluation as a clinician-in-the-loop tool.
发表机构
- Emory University School of Medicine(埃默里大学医学院)
- The University of Chicago(芝加哥大学)
- Winship Cancer Institute(温希普癌症研究所)
机构由 AI 辅助整理,请以论文原文为准。