arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13806cs.CLcs.AIcs.LG

理解智能体ICD编码的局限性

Understanding the Limits of Agentic ICD Coding

Chong Yock Eng, Yushi Cao, Yiming Chen, Kezhi Mao, Hongchao Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估神经、工作流和智能体系统在稀有度分层的MIMIC-IV数据上的ICD-10-CM编码性能,发现神经分类器在稀有代码上存在0.43微F1差距,工作流系统在损伤代码上近乎失败,而工具增强的智能体配置可恢复0.34微F1,但无单一系统全面占优。

中文摘要 AI 辅助

ICD-10-CM代码是用于美国医疗账单和流行病学报告中分类诊断和损伤的字母数字代码。标准的ICD-10-CM基准报告汇总指标,这些指标掩盖了复杂编码场景下的性能。我们在按稀有度分层的MIMIC-IV出院小结集合上评估了神经、工作流和智能体系统,并识别出两种正交的失败模式。神经分类器在稀有代码和常见代码之间表现出0.43的微F1差距。工作流系统处理稀有代码表现良好,但在需要多步指南遵循的损伤和外部原因代码上得分接近零。一种工具增强的智能体配置,能够结构化访问官方ICD-10-CM参考资料,在该子集上恢复了高达0.34的微F1。没有任何单一系统在所有条件下占主导地位。

英文摘要

ICD-10-CM codes are alphanumeric codes used in the US to classify diagnoses and injuries for medical billing and epidemiological reporting. Standard ICD-10-CM benchmarks report aggregate metrics that obscure performance on complex coding scenarios. We evaluate neural, workflow, and agentic systems on a rarity-stratified set of MIMIC-IV discharge summaries and identify two orthogonal failure modes. Neural classifiers exhibit a 0.43 micro-F1 gap between rare and common codes. Workflow systems handle rare codes well but score near zero on injury and external cause codes that require multi-step guideline following. A tool-augmented agentic configuration with structured access to official ICD-10-CM reference materials recovers up to 0.34 micro-F1 on this subset. No single system dominates across all conditions.

发表机构

  • Nanyang Technological University(南洋理工大学)
  • ASUS Intelligent Cloud Services (AICS)(华硕智能云服务)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑