arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10641cs.LGcs.AI

面向诊断预测的医学标记覆盖感知推理

Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction

  • Fuzhou University(福州大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

Kaisong Zhang, Haotian Fang, Junmeng Zhou, Hang Lv, Yulan Pan, Yanchao Tan

AI总结:

提出CARing框架,用组合语义ID表示诊断,通过覆盖奖励等改进多标签预测,在MIMIC-III、MIMIC-IV上的加权F1和top-k召回率优于基线。

AI中文摘要:

大型语言模型(LLM)凭借整合纵向临床证据并以自然语言对其进行推理的能力,为下次就诊诊断预测提供了广阔潜力。然而,针对LLM推理的强化学习通常根据最终答案的正确性对每个轨迹进行奖励。在下次就诊诊断预测中,多种诊断可同时有效,但每个轨迹仅奖励一种诊断的做法无法区分重复命中与不同诊断的覆盖情况,因此策略可能会集中于少数正确诊断,而遗漏其他诊断。同时,LLM的分词器会将ICD代码拆分为多个临床意义有限的通用标记,需多个解码步骤才能预测每种诊断,阻碍了对大型疾病词汇表的推理。为应对这两项挑战,我们提出CARing框架,该框架用组合语义ID(SID)表示诊断,并针对多标签覆盖优化推理轨迹。具体而言,我们首先通过残差量化将本体丰富的疾病语义编码为紧凑的SID,再通过多任务对齐和推理增强训练,将生成的SID标记与自然语言及纵向电子健康记录(EHR)上下文关联,以释放可迁移的LLM推理能力。CARing还通过强化学习的覆盖奖励和多正样本监督改进无序多标签预测。推理时,模型支持高效的直接约束解码和带排名融合的多链推理。在MIMIC-III和MIMIC-IV数据集上,CARing在加权F1指标上超过所有EHR训练的基线,且在所有报告的截断点均达到最高的top-k召回率,包括推理模式下的R@30分别为46.04%和46.52%。我们的代码和日志可在该https URL获取。

英文摘要:

Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language. However, reinforcement learning for LLM reasoning commonly rewards each trajectory according to the correctness of its final answer. In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses. The policy can therefore concentrate on a few correct diagnoses, leaving others uncovered. Meanwhile, LLM tokenizers can split ICD codes into several generic tokens with limited clinical meaning, requiring multiple decoding steps to predict each diagnosis and hindering reasoning over a large disease vocabulary. To address both challenges, we propose CARing, a framework that represents diagnoses with compositional Semantic IDs (SIDs) and optimizes reasoning trajectories for multi-label coverage. Concretely, we first encode ontology-enriched disease semantics into compact SIDs through residual quantization, and ground the resulting SID tokens in natural language and longitudinal EHR contexts through multi-task alignment and reasoning-enriched training to unlock transferable LLM reasoning. CARing further improves unordered multi-label prediction through a coverage reward for reinforcement learning and multi-positive supervision. At inference time, the model supports both efficient direct constrained decoding and multi-chain reasoning with rank fusion. On MIMIC-III and MIMIC-IV, CARing exceeds all EHR-trained baselines in weighted F1 and attains the highest top-k recall at every reported cutoff, including R@30 of 46.04% and 46.52% in reasoning mode. Our codes and logs are available at https://github.com/zmlxzyh/CARing-Codes-Logs.

↑