通过语义感知多模态预训练释放医疗表格数据的力量
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
浏览论文内容
中文总结 AI 辅助
该研究针对医疗表格数据的语义利用不足问题,提出语义感知多模态预训练框架,含重要性感知自适应掩码与软标签离散化模块,在皮肤病学、眼科数据集上实现SOTA,鲁棒性与跨域泛化能力优异。
中文摘要 AI 辅助
尽管视觉-语言模型在医疗表征学习中占据主导地位,但非结构化文本缺乏结构化临床表格所固有的密集、定量诊断表型。然而,现有的多模态预训练方法由于采用与语义无关的设计(将表格输入视为扁平向量)以及不稳定的连续回归目标,未能充分利用这一潜力。为克服该问题,我们提出一种新颖的语义感知框架,明确建模表格数据的内在二维结构:其一,针对特征间具有不同诊断重要性的层级结构,引入重要性感知自适应掩码,构建无标签课程以优先考虑显著特征;其二,针对特征内的连续性-离散性二元性,提出软标签离散化模块,用稳定的分布匹配替代不稳定的数值回归,从而在数学上保留序数关系。在大规模皮肤病学数据集(SLICE-3D、HOP)和眼科数据集(EyePACS)上开展的大量实验建立了新的 state-of-the-art(SOTA),展现出卓越的鲁棒性和跨域泛化能力。
英文摘要
While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables. However, existing multimodal pre-training methods underutilize this potential due to semantic-agnostic designs that treat tabular inputs as flat vectors and employ unstable continuous regression objectives. To overcome this, we propose a novel semantic-aware framework explicitly modeling the intrinsic two-dimensional structure of tabular data. First, addressing the inter-feature hierarchy of varying diagnostic importance, we introduce Importance-Aware Adaptive Masking to construct a label-free curriculum prioritizing salient features. Second, addressing the intra-feature continuity-discreteness duality, we propose a Soft-Label Discretized Module that replaces unstable numerical regression with stable distribution matching, thereby mathematically preserving ordinal relationships. Extensive experiments across large-scale dermatology (SLICE-3D, HOP) and ophthalmology (EyePACS) datasets establish a new state-of-the-art (SOTA), demonstrating exceptional robustness and cross-domain generalizability.
发表机构
- Faculty of Information Technology, Monash University(莫纳什大学信息技术学院)
- Monash University(莫纳什大学)
- The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。