arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17855cs.AI

SNOMED CT 概念推荐:基于掩码临床上下文

SNOMED CT Concept Recommendation from Masked Clinical Context

Ali Noori

首次发表
浏览论文内容

中文总结 AI 辅助

针对临床概念推荐中罕见概念难题,基于掩码上下文基准,对比多种方法,发现稀疏TF-IDF最优,且概念频率显著影响性能,为低资源场景提供可复现基线。

中文摘要 AI 辅助

将临床语言标准化为 SNOMED CT 可支持互操作性、分析和可复用的表型分析,但当相关概念在训练数据中罕见或缺失时,概念推荐仍然困难。我们提出了一个掩码概念推荐基准,使用源自 MIMIC-IV-Note 的 SNOMED CT 实体链接挑战 v1.2.1 数据。该数据集包含 272 份出院小结中的 75,491 条标注,其中 204 份笔记用于训练,68 份用于历史测试。对于每个唯一的笔记-概念对,目标提及从局部临床上下文中被掩码,系统对训练期间观察到的 SNOMED CT 概念进行排序。我们比较了流行度基线、稀疏 TF-IDF 概念原型、密集潜在语义分析嵌入、稀疏-密集融合、检索笔记证据以及检索增强混合方法。稀疏 TF-IDF 表现最佳,达到 Recall@1 14.81%、Recall@10 33.43%、MRR 0.2114 和 nDCG@10 0.2297。检索增强未改善该基线,Recall@10 为 31.99%,MRR 为 0.1937。性能受概念频率强烈影响:仅出现在一或两份训练笔记中的概念 Recall@10 为 7.74%,而出现在超过十份笔记中的概念为 43.90%。此外,9.66% 的测试笔记-概念对包含训练期间未见的概念。这些发现表明,局部词汇上下文和术语覆盖率是低资源环境下推荐质量的主要决定因素,并为未来基于本体和生物医学编码器的检索系统提供了可复现的基线。

英文摘要

Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant concepts are rare or absent from training data. We present a masked-concept recommendation benchmark using the SNOMED CT Entity Linking Challenge v1.2.1 data derived from MIMIC-IV-Note. The dataset contains 75,491 annotations across 272 discharge summaries, with 204 notes used for training and 68 for historical testing. For each unique note-concept pair, the target mention is masked from a local clinical context and the system ranks SNOMED CT concepts observed during training. We compare a popularity baseline, sparse TF-IDF concept prototypes, dense latent semantic analysis embeddings, sparse-dense fusion, retrieved-note evidence, and a retrieval-augmented hybrid. Sparse TF-IDF performs best, achieving Recall@1 of 14.81%, Recall@10 of 33.43%, MRR of 0.2114, and nDCG@10 of 0.2297. Retrieval augmentation does not improve this baseline, with Recall@10 of 31.99% and MRR of 0.1937. Performance is strongly affected by concept frequency: Recall@10 is 7.74% for concepts appearing in only one or two training notes versus 43.90% for concepts appearing in more than ten. In addition, 9.66% of test note-concept pairs contain concepts unseen during training. These findings show that local lexical context and terminology coverage are major determinants of recommendation quality in low-resource settings and provide a reproducible baseline for future ontology-grounded and biomedical-encoder retrieval systems.

发表机构

  • University of North Carolina Greensboro(北卡罗来纳大学格林斯伯勒分校)

机构由 AI 辅助整理,请以论文原文为准。

↑