arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.09699cs.CLcs.AI

结构化信息很重要:基于患者级知识图谱的可解释ICD编码

Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs

  • The University of Manchester(曼彻斯特大学)
  • Imperial College London(伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

Mingyang Li, Viktor Schlegel, Tingting Mu, Warren Del-Pinto, Goran Nenadic

更新

AI总结:

该研究提出利用文档级知识图谱构建输入文档的结构化表示,将其集成至PLM-ICD架构中进行自动化ICD编码,在提升Macro-F1分数与训练效率的同时增强了可解释性。

AI中文摘要:

将临床文档映射到标准化临床词汇表是一项重要任务,因为它为信息检索和分析提供结构化数据,这对于临床研究、医院管理和改善患者护理至关重要。然而,手动编码既困难又耗时,使其在大规模下不切实际。自动化编码可以潜在地减轻这一负担,提高结构化临床数据的可用性和准确性。该任务难以自动化,因为它需要映射到高维和长尾目标空间,例如国际疾病分类(ICD)。虽然外部知识源已被广泛用于增强输出代码表示,但外部资源在表示输入文档方面的使用却未得到充分探索。在这项工作中,我们计算输入文档的结构化表示,利用文档级知识图谱(KGs)提供患者病情的全面结构化视图。由此产生的知识图谱以原始文本的23%高效表示以患者为中心的输入文档,同时保留90%的信息。我们通过将其集成到最先进的ICD编码架构PLM-ICD中,评估了该图谱在自动化ICD-9编码中的有效性。我们的实验在流行的基准测试上使Macro-F1分数提高了多达3.20%,同时提高了训练效率。我们将这一改进归因于KG中不同类型的实体和关系,并展示了该方法相比纯文本基线在可解释性方面的提升潜力。

英文摘要:

Mapping clinical documents to standardised clinical vocabularies is an important task, as it provides structured data for information retrieval and analysis, which is essential to clinical research, hospital administration and improving patient care. However, manual coding is both difficult and time-consuming, making it impractical at scale. Automated coding can potentially alleviate this burden, improving the availability and accuracy of structured clinical data. The task is difficult to automate, as it requires mapping to high-dimensional and long-tailed target spaces, such as the International Classification of Diseases (ICD). While external knowledge sources have been readily utilised to enhance output code representation, the use of external resources for representing the input documents has been underexplored. In this work, we compute a structured representation of the input documents, making use of document-level knowledge graphs (KGs) that provide a comprehensive structured view of a patient's condition. The resulting knowledge graph efficiently represents the patient-centred input documents with 23\% of the original text while retaining 90\% of the information. We assess the effectiveness of this graph for automated ICD-9 coding by integrating it into the state-of-the-art ICD coding architecture PLM-ICD. Our experiments yield improved Macro-F1 scores by up to 3.20\% on popular benchmarks, while improving training efficiency. We attribute this improvement to different types of entities and relationships in the KG, and demonstrate the improved explainability potential of the approach over the text-only baseline.

↑