发表机构
Bangladesh University of Engineering and Technology(孟加拉工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对孟加拉国X光片骨折分类,提出可靠性与解剖一致性感知的多模态融合方法,在提升分类性能的同时降低元数据不匹配带来的损失,存在干净性能与鲁棒性的权衡。
AI 中文摘要
背景:多模态骨折分类器可从患者及解剖元数据中获益,但当上下文信息缺失或不匹配时,其性能会变得脆弱。方法:我们使用孟加拉国OrthoFrac-XR数据集的1493张X光片,采用无信息泄露的年龄、性别、骨类型和侧别信息,将ConvNeXt图像编码器与临床多层感知器通过拼接、晚期融合、可靠性门控残差融合及分层状态-位置公式结合,还引入了解剖一致性门:当图像侧解剖预测与报告的骨类型不一致时,减弱元数据修正。结果:在5折交叉验证和3个随机种子下,分层残差融合的宏F1值为0.6046±0.0279,仅图像学习的宏F1值为0.5727±0.0270,同时Brier评分从0.5239降至0.4948;在5折鲁棒性实验中,与普通残差融合相比,解剖一致性融合在元数据打乱时将宏F1损失从0.0567降至0.0203,不过其干净数据下的宏F1更低;推理时无骨类型信息,辅助解剖监督使宏F1从0.5620±0.0330提升至0.5899±0.0289。结论:结构化上下文可提升骨折分类性能,一致性感知门控可限制元数据不匹配的危害,观察到的干净性能与鲁棒性的权衡及无患者级标识符的情况,推动了外部和前瞻性验证。
英文摘要
Background: Multimodal fracture classifiers may benefit from patient and anatomical metadata, but they can also become brittle when contextual information is missing or mismatched. Methods: We studied 1493 radiographs from the Bangladeshi OrthoFrac-XR dataset using leakage-safe age, sex, bone type, and laterality. A ConvNeXt image encoder was combined with a clinical multilayer perceptron through concatenation, late fusion, reliability-gated residual fusion, and a hierarchical state-location formulation. We additionally introduced an anatomy-consistency gate that attenuates metadata corrections when an image-side anatomical prediction disagrees with the reported bone type. Results: Across five folds and three seeds, hierarchical residual fusion achieved a macro-F1 of 0.6046 +/- 0.0279, compared with 0.5727 +/- 0.0270 for image-only learning, while improving the Brier score from 0.5239 to 0.4948. In a five-fold robustness experiment, anatomy-consistency fusion reduced the macro-F1 loss under shuffled metadata from 0.0567 to 0.0203 relative to ordinary residual fusion, although its clean-data macro-F1 was lower. Without bone type at inference, auxiliary anatomy supervision improved macro-F1 from 0.5620 +/- 0.0330 to 0.5899 +/- 0.0289. Conclusions: Structured context improves fracture classification, and consistency-aware gating limits harm from mismatched metadata. The observed clean-performance-robustness trade-off and the absence of patient-level identifiers motivate external and prospective validation.