arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FiberGeoText:用于浅层白质群体级组织的视觉-语言模型

FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter

Yuqian Chen, R. Jarrett Rushmore, Guikun Chen, Fan Zhang, Edward Yeterian, Nikos Makris, Yogesh Rathi, Lauren J. O'Donnell

arXiv 2610.02755首次发表:更新:

发表机构

Mass General Brigham; Harvard Medical School; Boston University Chobanian & Avedisian School of Medicine; Colby College(麻省总医院布莱根医疗系统; 哈佛医学院; 波士顿大学乔巴尼安与阿维迪西安医学院; 科尔比学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对浅层白质短程纤维组织难以跨个体对应的问题,提出视觉-语言模型FiberGeoText,融合轨迹、皮层解剖和形状信息进行群体级聚类,在0.76 mm扩散MRI数据上优于SOTA,并泛化至未见受试者。

AI 中文摘要

浅层白质(SWM)是贯穿生命全程的认知和脑疾病的关键脑区,包含丰富的短程联合纤维,其组织尚未被完全表征,部分原因在于短轨迹和高度可变的皮层折叠使得个体间的对应关系具有挑战性。解剖学上对应的连接在个体间可能具有不同的空间位置,因此仅凭几何邻近性可能无法充分定义。我们引入了FiberGeoText(FGT),一种视觉-语言模型(VLM),用于将从超高分辨率扩散MRI重建的短程浅层白质(SWM)流线组织成群体级聚类。FGT联合表示每条流线的三个互补属性:其三维轨迹、皮层解剖背景和形状。来自多个分区方案的皮层端点信息以文本形式表达,并使用预训练的大语言模型(LLM)进行编码,使得异质解剖描述能够贡献于共同的连续表示。我们在采集的亚毫米0.76 mm扩散MRI数据上评估了FGT。与最先进(SOTA)方法相比,FGT产生了显著更高的皮层分区一致性、簇内形状一致性、簇大小一致性和跨受试者对应性。训练后的模型也能很好地泛化到未见过的受试者,平均恢复了5,000个学习簇中的96.7%,并且训练和测试数据之间的簇结构具有高度一致性。总之,这些发现表明,通过使用VLM模型学习多模态深度嵌入来整合几何、解剖和形状信息,能够在个体间解剖变异的情况下实现稳健的群体一致性SWM组织学习。

英文摘要

The superficial white matter (SWM), a critical brain region for cognition across the lifespan and brain disease, contains abundant short-range association fibers whose organization remains incompletely characterized, in part because the short trajectories and highly variable cortical folding make correspondence across individuals challenging. Anatomically corresponding connections may vary in spatial location across individuals and therefore may not be adequately defined by geometric proximity alone. We introduce FiberGeoText (FGT), a vision-language model (VLM) for organizing short-range superficial white matter (SWM) streamlines reconstructed from ultra-high-resolution diffusion MRI into population-level clusters. FGT jointly represents three complementary properties of each streamline: its three-dimensional trajectory, its cortical anatomical context, and its shape. Cortical endpoint information from multiple parcellation schemes is expressed as text and encoded using a pretrained large language model (LLM), enabling heterogeneous anatomical descriptions to contribute to a common continuous representation. We evaluated FGT on acquired submillimeter 0.76 mm diffusion MRI data. Compared with state-of-the-art (SOTA) methods, FGT produced substantially greater cortical parcel coherence, within-cluster shape consistency, cluster-size consistency, and cross-subject correspondence. The trained model also generalizes well to unseen subjects with an average of 96.7% of the 5,000 learned clusters recovered, and high consistency of cluster structure between training and testing data. Together, these findings demonstrate that integrating geometric, anatomical, and shape information by learning multimodal deep embeddings with a VLM model enables robust learning of population-consistent SWM organization despite interindividual anatomical variability.

Comments22 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑