arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19721cs.AIcs.MA

LearnActCoder:用于自适应临床编码智能体的角色感知错误记忆

LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents

  • Optum AI(奥普特姆人工智能)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

Meysam Ghaffari, Bhaskar Sen, Nasim Sabetpour, Nina Fatehi, Animesh Agarwal, Carlos Morato

AI总结:

提出LearnActCoder,利用结构化错误记忆(MistakeKDB)在推理时自适应调整临床编码智能体,通过编码器-评判器分工提升CPT F1达5.9个百分点,且无需权重更新。

AI中文摘要:

临床编码智能体反复遇到相同的失败模式,包括不支持的编码、遗漏的记录病症、特异性错误以及手术编码约定不匹配。我们引入了“先学后行动”(Learn-Then-Act),一种推理时自适应框架,将来自小型标注LEARN批次中的错误转换为结构化的错误知识数据库(MistakeKDB)。假阴性教训被路由到面向回忆的编码器(Coder),而假阳性教训被路由到面向精确度的评判器(Judge)。我们将该框架实例化为LearnActCoder,一个带有查找表接地(在可用时)的编码器-评判器临床编码流水线。在150份匹配的MIMIC-III病历上,结构化的MistakeKDB将CPT F1提高了5.9个百分点,而原始示例和反思式记忆仍接近无记忆基线;ICD-9的改进不显著。在匹配的MIMIC-IV队列中,记忆使ICD-10编码转向更高的精确度,但以召回率为代价,F1在统计上保持不变。将相同的记忆应用于1,000份保留的MIMIC-III病历,保持了稳定的ICD工作点,提供了规模/稳定性证据。总体而言,结果与结构化、反馈衍生的错误记忆有助于在无需权重更新或改变底层工作流程的情况下跨病例调整临床编码行为这一观点一致。绝对CPT/HCPCS性能仍然较低,且系统是回顾性评估的,而非在临床部署中评估。

英文摘要:

Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introduce Learn-Then-Act, an inference-time adaptation framework that converts errors from a small labeled LEARN batch into a structured Mistake Knowledge Database (MistakeKDB). False-negative lessons are routed to a recall-oriented Coder, while false-positive lessons are routed to a precision-oriented Judge. We instantiate the framework in LearnActCoder, a Coder-Judge clinical coding pipeline with lookup-table grounding where available. On 150 matched MIMIC-III notes, structured MistakeKDB improves CPT F1 by 5.9 percentage points, while raw-example and reflection-style memories remain near the no-memory baseline; the ICD-9 improvement is not significant. On a matched MIMIC-IV cohort, memory shifts ICD-10 coding toward higher precision at a recall cost, leaving F1 statistically unchanged. Applying the same memory to 1,000 held-out MIMIC-III notes maintains a stable ICD operating point, providing scale/stability evidence. Overall, the results are consistent with structured, feedback-derived error memory being useful for adapting clinical coding behavior across cases without weight updates or changes to the underlying workflow. Absolute CPT/HCPCS performance remains low, and the system is evaluated retrospectively rather than in clinical deployment.

↑