发表机构
Xinjiang University; Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center; Tsinghua University(新疆大学; 新疆多模态智能处理与信息安全工程技术研究中心; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对放射学报告生成中视觉异常线索易被文本先验稀释的问题,提出同质语义对齐与层级专家路由框架,通过消除跨模态分布偏移和动态专家路由,在三个基准数据集上达到最优性能。
AI 中文摘要
放射学报告生成(RRG)旨在将医学图像转换为诊断文本,以辅助临床决策并减轻医生的工作负担。尽管现有方法在跨模态交互和外部先验知识融入方面取得了广泛进展,但底层表征的分布偏移以及异构信息的无差别刚性耦合,导致在解码过程中,微弱的视觉异常线索容易被大量的文本先验和生成惯性所稀释。为克服这一瓶颈,受认知科学启发,我们提出了一种新颖的两阶段同质语义对齐与层级专家路由(HSA-HER)框架。首先,该模型在底层潜在空间中引入显式的同质分布约束,有效消除视觉特征与文本特征之间的跨模态分布偏移,从而提取出纯净的视觉特征作为与疾病精准对齐的语义锚点。其次,针对由视觉特征、局部实体和全局检索构成的异构临床证据,我们设计了一种由这些疾病语义锚点引导的层级专家路由机制。该机制摒弃了无差别的刚性耦合范式,具体而言,它动态激活专家网络,对多源证据进行定向挖掘和语义重构,并自适应地分配融合权重。在三个主流基准数据集上进行的大量实验表明,HSA-HER取得了最先进的性能,能够准确描绘复杂的影像细节和关键诊断信息。
英文摘要
Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of underlying representations and the undifferentiated rigid coupling of heterogeneous information cause weak visual abnormality cues to be easily diluted by massive text priors and generation inertia during decoding. To overcome this bottleneck, inspired by cognitive science, we propose a novel two-stage Homogeneous Semantic Alignment and Hierarchical Expert Routing (HSA-HER) framework. First, the model introduces an explicit homogeneous distribution constraint in the underlying latent space to effectively eliminate the cross-modal distribution shift between visual and textual features, thereby extracting purified visual features as semantic anchors that accurately align with diseases. Second, for heterogeneous clinical evidence composed of visual features, local entities, and global retrievals, we design a hierarchical expert routing mechanism guided by these disease semantic anchors. This mechanism abandons the undifferentiated rigid coupling paradigm. Specifically, it dynamically activates expert networks to perform targeted mining and semantic reconstruction on multi-source evidence, and adaptively allocates fusion weights. Extensive experiments on three mainstream benchmark datasets demonstrate that HSA-HER achieves state-of-the-art performance, accurately depicting complex imaging details and key diagnostic information.