发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将自动化临床编码中的系统性差异建模为编码风格,证明条件化于风格可显著提升ICD编码F1分数,揭示单金标准评估中的误差部分实为可恢复的风格因素。
AI 中文摘要
在自动化临床编码中,标签空间涵盖数万个诊断和程序代码,模型目前针对单一金标准标注进行评估,将任何偏差视为误差。但我们发现,当两个团队对相同的110个ACI-Bench病例进行编码时,他们对同一份病历的代码一致性仅为73%(Jaccard相似度);即使在独立临床审计移除错误代码后,一致性也仅上升至77%。这一差距是误差还是系统性因素?我们将系统性成分建模为编码风格ψ,即编码员或站点特定的关于编码内容及记录程度的策略,并将编码重新表述为p(代码|病历,ψ),使用10维评分标准估计ψ。如果风格是噪声,条件化于它不会产生任何效果。相反,在五个数据集中,使用数据匹配风格的条件化模型将ICD F1分数最高提升26点,而极端不匹配风格则使其降低最多21点。四种基于提示的编码方法在39-49 F1范围内,一旦提供风格后收敛至52-56(所有p<0.05)。单一金标准评估归咎于模型误差的很大一部分是可恢复的、未被建模的风格。
英文摘要
In automated clinical coding, where the label space spans tens of thousands of diagnosis and procedure codes, models are currently evaluated against a single gold annotation, treating any deviation as error. But we find when two teams code the same 110 ACI-Bench encounters, they agree on only 73% of codes (Jaccard similarity) for the same note; even after an independent clinical audit removes erroneous codes, agreement rises only to 77%. Is that gap error or something systematic? We model the systematic component as coding style $ψ$, a coder- or site-specific policy over what to code and how much to document, and recast coding as $p(\mathrm{code}\mid\mathrm{note},ψ)$, estimating $ψ$ with a 10-dimension rubric. If style were noise, conditioning on it would do nothing. Instead, across five datasets a model conditioned with a data-matching style raises ICD F1 by up to 26 points and an extreme mismatched one lowers it by up to 21. Four prompt based coding methods spanning 39-49 F1 converge to 52-56 once style is supplied (All p<0.05). Much of what single-gold evaluation charges to model error is recoverable, unmodeled style.