arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解剖感知的细粒度多模态融合用于基于CT和放射学报告的喉咽癌T分期预测

Anatomy-aware Fine-grained Multimodal Fusion for Laryngopharyngeal Cancer T-Staging Prediction Using CT and Radiology Report

Xingyue Zhao, Yanzhou Su, Fang Zhang, Zhanghexuan Ji, Yirui Wang, Dazhou Guo, Sibo Ju, Yuehua Cheng, Yuzhen Chen, Ming Feng, Le Lu, Tsung-Ying Ho, Jian Wang, Dakai Jin, Na Shen

arXiv 2610.06837首次发表:更新:

发表机构

DAMO Academy, Alibaba Group; Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College; Hupan Laboratory; Fuzhou University; Eye & ENT Hospital, Fudan University; Zhongshan Hospital, Fudan University; Chang Gung Memorial Hospital; Zhongshan Hospital(Xiamen), Fudan University(阿里巴巴集团达摩院; 中国医学科学院北京协和医学院北京协和医院; 湖畔实验室; 福州大学; 复旦大学附属眼耳鼻喉科医院; 复旦大学附属中山医院; 长庚纪念医院; 复旦大学附属中山医院厦门医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对喉咽癌T分期中结构建模与跨模态对齐不足,提出解剖感知多模态框架,通过器官图构建、跨模态对齐与报告增强细化,实现优越性能。

AI 中文摘要

准确的T分期对于指导喉咽癌的个性化治疗策略至关重要。然而,当前的临床实践依赖于侵入性活检程序,而基于CT的分期由于肿瘤侵袭的复杂模式仍然具有挑战性。最近的计算机辅助方法面临两个关键挑战:1)结构关系建模:现有方法未能充分表征肿瘤侵袭的解剖结构化模式,因为它们要么在没有肿瘤特异性解剖约束的情况下处理整个CT体积,要么依赖劳动密集型的肿瘤分割。2)细粒度跨模态对齐:虽然放射学报告包含器官特异性的侵袭细节,但当前应用全局特征融合的方法难以准确地将个体解剖结构与相应的文本描述对齐。为解决这些问题,我们提出了一种解剖感知的多模态框架,将器官级CT上下文和放射学报告整合为喉咽T分期的统一表示。该框架首先构建一个解剖结构化器官图(AOG),捕获原发部位与周围器官之间的侵袭模式,然后执行器官锚定跨模态对齐(OCA),使每个器官节点聚合来自放射学报告的文字证据,最后通过报告增强图细化(REG)注入从报告中提取的器官特异性侵袭线索来细化该图表示,产生结合空间和文字证据的多模态器官图。大量实验表明,所提出的框架在喉咽癌T分期中实现了优越的性能。

英文摘要

Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.

CommentsAccepted by IEEE Transactions on Medical Imaging (IEEE TMI)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑