arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30234cs.AI

CoLa-ICD:一种用于长尾自动医学编码的知识增强框架

CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding

  • University of Otago(奥塔哥大学)

机构由 AI 辅助整理,请以论文原文为准。

Yihang Cheng, Veronica Liesaputra, Andrew Trotman

AI总结:

针对自动医学编码中罕见ICD编码的长尾挑战,本文提出知识增强框架CoLa-ICD,通过丰富标签、建模编码依赖及对齐语义与临床证据,在AUC、F1等指标上实现最先进性能。

AI中文摘要:

自动医学编码是指将ICD编码分配给临床病历,但由于病历冗长、标签分布不均衡以及术语多样,该任务仍具挑战性,这些挑战对于训练实例有限、易与语义相似标签混淆的罕见编码尤为严重。本文提出CoLa-ICD,一种用于长尾预测的知识增强框架,该框架通过外部术语丰富ICD标签、建模相关编码间的依赖关系,并学习标签语义与临床证据间更强的对齐以实现长尾预测。实验表明,CoLa-ICD在更大、更稀疏的标签空间中长尾预测提升幅度更大,在AUC、F1和P@k指标上达到了最先进性能,其代码可在指定URL获取。

英文摘要:

Automatic medical coding assigns ICD codes to clinical notes, but it remains challenging due to long documents, imbalanced label distributions, and diverse terms. These challenges are especially severe for rare codes, which have limited training instances and are easily confused with semantically similar labels. We introduce CoLa-ICD, a knowledge-enhanced framework for long-tail prediction. CoLa-ICD enriches ICD labels with external terms, models dependencies among related codes, and learns stronger alignment between label semantics and clinical evidence for long-tail prediction. Experiments show that CoLa-ICD improves long-tail prediction with larger gains in larger and sparser label spaces and achieves state-of-the-art performance in AUC, F1, and P@k. Our code is available at https://github.com/youwillbethebest/Cola-ICD.

↑