arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于极端临床代码预测的图约束策略学习

Graph-Constrained Policy Learning for Extreme Clinical Code Prediction

Amritpal Singh, Sebastian Torres, Khawar Shakeel, Syed Ahmad Chan Bukhari

arXiv 2607.11954首次发表:更新:

AI 中文总结

研究针对临床代码预测任务,提出图约束遍历策略,将其转换为有限时域决策过程。单一语言模型逐层选节点得叶代码,实现极端多标签预测到子集决策的转换。实验表明该策略优于平面基线,增加监督数据可提升性能,简单图约束策略学习效果良好。

AI 中文摘要

临床代码预测将非结构化出院小结映射到大型、稀疏且深度分层标签空间中的ICD - 10 - CM叶代码。多数系统将此任务视为平面多标签分类,对代码独立评分,为稀有标签提供的训练信号有限。我们提出图约束遍历策略,将ICD预测制定为在修剪后的代码层次结构上的有限时域决策过程。单一语言模型逐层下降图,选择有效子节点直至到达可计费叶代码。这将极端多标签预测转换为稀疏、层次感知的子集决策,同时保证结构有效的输出。在MIMIC - IV出院小结上,我们最佳的监督策略SFT - 1 +在精心策划的50代码子集上实现了0.709的微F1,在完整的15,761代码空间上实现了0.527的微F1,优于包括CAML、LAAT和PLM - ICD在内的平面基线。在完整设置中,SFT - 1 +比最强的平面基线在微F1上提高了0.044,在宏F1上提高了0.157,表明图约束分解减轻了稀有代码瓶颈。一项受控析因研究评估了架构、训练算法和数据预算。在两个规模上,一个共享策略匹配三位专家级联,同时在28 - 32%的全空间测试记录上避免其上下文窗口溢出。增加监督轨迹数据是唯一持续提高性能的干预措施,而GRPO强化学习与匹配数据的监督延续相比没有优势。这些结果表明,简单的图约束策略学习在极端临床代码预测方面可以优于更复杂的平面、级联和强化学习替代方案。

英文摘要

Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label space. Most systems treat the task as flat multi-label classification, scoring codes independently and providing limited training signal for rare labels. We propose a graph-constrained traversal policy that formulates ICD prediction as a finite-horizon decision process over a pruned code hierarchy. A single language model descends the graph level by level, selecting valid child nodes until billable leaf codes are reached. This converts extreme multi-label prediction into sparse, hierarchy-aware subset decisions while guaranteeing structurally valid outputs. On MIMIC-IV discharge summaries, our best supervised policy, SFT-1+, achieves 0.709 micro-F1 on a curated 50-code subset and 0.527 micro-F1 on the full 15,761-code space, outperforming flat baselines including CAML, LAAT, and PLM-ICD. In the full setting, SFT-1+ improves over the strongest flat baseline by 0.044 micro-F1 and 0.157 macro-F1, suggesting that graph-constrained decomposition mitigates the rare-code bottleneck. A controlled factorial study evaluates architecture, training algorithm, and data budget. Across both scales, one shared policy matches a three-specialist cascade while avoiding its context-window overflow on 28-32% of full-space test notes. Increasing supervised trajectory data is the only intervention that consistently improves performance, while GRPO reinforcement learning provides no benefit over supervised continuation with matched data. These results show that simple graph-constrained policy learning can outperform more complex flat, cascaded, and reinforcement-learning alternatives for extreme clinical code prediction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑