arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REFINE:面向带预算的文本属性图的大语言模型优化,用于个性化医学概念表示

REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation

Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Dongjie Wang, Zijun Yao

arXiv 2609.04415首次发表:更新:

发表机构

University of Kansas(堪萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对电子健康记录预测中患者个性化医学概念表示的需求,提出REFINE框架,结合强化学习、异构GNN与冻结LLM,在MIMIC-III和MIMIC-IV数据集上提升了EHR预测性能。

AI 中文摘要

学习丰富的医学概念表示对于电子健康记录(EHR)预测至关重要。文本属性知识图谱(TKG)通过将异构医学关系与文本语义组织在一起,提供了一个天然的基础。然而,大多数现有的编码器对所有患者的概念进行统一处理,尽管某一代码的含义和预测价值取决于患者特定的临床背景和轨迹。从TKG中学习患者个性化的概念表示带来了两个关键挑战:(1)决定为每个观察到的代码纳入多少知识图谱(KG)上下文;(2)使语义信息与患者特定的关系结构对齐。我们提出了REFINE,一个用于患者个性化医学概念编码的感知KG的带预算大语言模型(LLM)图优化框架。从全局TKG开始,REFINE构建患者特定的时间图。一个顺序强化学习策略为每个观察到的代码选择个性化的KG扩展预算。得到的患者图由异构图神经网络(GNN)处理,以捕获感知关系的结构依赖关系,而冻结的LLM使用感知图的软提示来语义优化概念表示。在MIMIC-III和MIMIC-IV上的实验表明,REFINE能持续改进各种EHR主干模型,优于强大的基线模型,并且在组件消融、KG选择和数据不足的情况下表现出稳健的增益。

英文摘要

Learning rich medical concept representations is essential for EHR prediction. Text-attributed knowledge graphs (TKGs) provide a natural foundation by organizing heterogeneous medical relations together with textual semantics. However, most existing encoders process concepts uniformly across patients, despite the fact that a code's meaning and predictive value depend on patient-specific clinical context and trajectory. Learning patient-personalized concept representations from TKGs introduces two key challenges: (1) deciding how much KG context to incorporate for each observed code, and (2) aligning semantic information with the patient-specific relational structure. We propose REFINE, a KG-aware budgeted LLM graph refinement framework for patient-personalized medical concept encoding. Starting from a global TKG, REFINE constructs patient-specific temporal graphs. A sequential reinforcement learning policy selects a personalized KG expansion budget for each observed code. The resulting patient graph is processed by a heterogeneous GNN to capture relation-aware structural dependencies, while a frozen LLM uses graph-aware soft prompts to semantically refine concept representations. Experiments on MIMIC-III and MIMIC-IV show that REFINE consistently improves diverse EHR backbones, outperforms strong baselines, and demonstrates robust gains across component ablation, KG selection, and data insufficiency.

CommentsThis paper has been accepted at the EMNLP 2026 main conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑