arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TractoGraphVLM:用于白质纤维束成像的统一视觉-语言框架

TractoGraphVLM: A Unified Vision-Language Framework for White Matter Tractography

Gurucharan Marthi Krishna Kumar, Janine Dale Mendola, Amir Shmuel

arXiv 2608.18166首次发表:更新:

发表机构

Montreal Neurological Institute; McGill University; Department of Ophthalmology, McGill University; McConnell Brain Imaging Centre(蒙特利尔神经学研究所; 麦吉尔大学; 麦吉尔大学眼科系; 麦康奈尔脑成像中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TractoGraphVLM是统一视觉-语言框架,可完成白质纤维束的分类、检索、描述、问答四项任务,在HCP数据集上表现良好,具跨年龄迁移鲁棒性,仅从语言学习神经解剖学知识。

AI 中文摘要

视觉语言模型已在二维医学成像领域取得突破,但将其扩展至三维白质纤维束成像仍面临挑战,原因在于纤维束的复杂拓扑结构。我们提出TractoGraphVLM,这是一个面向四项任务(纤维束分类、文本-纤维束检索、解剖学图像描述、视觉问答)的统一框架,基于共享的GPS架构、训练流程与读出设计构建。纤维束被表示为流线图,其节点编码三维位置与切线方向。通用、强大、可扩展(GPS)的图变换器生成与冻结的BiomedBERT文本编码器通过对比学习对齐的纤维束嵌入,而带有视觉前缀标记的BioGPT解码器则生成描述文本与答案。单个共享编码器与解码器在所有四项任务上联合训练,并从一个检查点进行评估。在HCP年轻成年受试者上训练后,TractoGraphVLM在保留测试集上取得91.8%的纤维束分类准确率、84.7%的检索R@1、BLEU-4=20.1、ROUGE-L=66.8以及66.4%的视觉问答准确率。相同检查点零样本迁移至HCP老年受试者时,判别式任务出现适度下降,生成式任务下降幅度更大,显示出对年龄与采集偏移的鲁棒性。语言监督比仅标签训练能产生更丰富的表示,恢复出半球、纤维家族等结构,这些结构由描述文本携带但从未作为标签提供。仅替换视觉编码器时,保留纤维方向的图优于体积基线,GPS取得最佳平衡。生成式指标衡量与结构化知识库的一致性而非独立临床文本;即便如此,TractoGraphVLM表明,对白质纤维束进行分类、检索、描述与问答可通过一个联合训练的模型实现,该模型仅从语言中学习可迁移的神经解剖学知识。

英文摘要

Vision language models have transformed 2D medical imaging, yet extending them to 3D white matter tractography remains challenging due to the complex topology of fiber bundles. We introduce TractoGraphVLM, a unified framework for four tasks, bundle classification, text-to-tract retrieval, anatomical captioning, and visual question answering, built on a shared GPS architecture, training procedure, and read-out design. Fiber bundles are represented as streamline graphs whose nodes encode 3D position and tangent orientation. A General, Powerful, Scalable (GPS) graph transformer produces bundle embeddings aligned with a frozen BiomedBERT text encoder via contrastive learning, while a BioGPT decoder with visual prefix tokens generates captions and answers. A single shared encoder and decoder is trained jointly across all four tasks and evaluated from one checkpoint. Trained on HCP Young Adult subjects, TractoGraphVLM achieves 91.8% bundle classification accuracy, 84.7% retrieval R@1, BLEU-4=20.1, ROUGE-L=66.8, and 66.4% VQA accuracy on a held-out test set. The same checkpoints transfer zero-shot to HCP Aging subjects, with a modest drop on discriminative tasks and a larger drop on generative tasks, showing robustness to age and acquisition shift. Language supervision yields richer representations than label-only training, recovering structure like hemisphere and fiber family, carried by captions but never given as a label. Swapping only the visual encoder, graphs preserving fiber orientation outperform volumetric baselines, with GPS giving the best balance. Generative metrics measure consistency with a structured knowledge base rather than independent clinical text; even so, TractoGraphVLM shows that classifying, retrieving, describing, and answering questions about a white matter bundle can be served by one jointly trained model that learns transferable neuroanatomy from language alone.

CommentsAccepted as a Spotlight at the ECCV 2026 Workshop on Artificial Intelligence for Medical 3D Vision (AI4M3D). Our codebase, including all training and evaluation pipelines, is publicly available at https://github.com/AS-Lab/Marthi-et-al-2026-TractoGraphVLM-Unified-Vision-Language-White-Matter-Tractography

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑