arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多病老年患者个性化用药安全的耦合图-策略蒸馏方法

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

Zihan Wang, Anglin Liu, Rongyi Wang, Dantong Li, Yi Lu, Siqing Yuan, Hongxia Xu, Zhongtian Long, Jintai Chen

arXiv 2608.09443首次发表:更新:

AI 中文总结

本文提出ATLAS框架与GeriMedBench基准,通过耦合图-策略蒸馏实现多病老年患者的个性化用药安全,在多基准测试中表现优于对比系统,获临床医生更高评分。

AI 中文摘要

大型语言模型(LLM)智能体可在临床就诊间隔期间支持药物审查,但多病老年患者的安全用药选择取决于患者可能遗漏的病情、药物及老年风险因素。本文提出ATLAS,这是一种用于患者自适应用药安全的耦合图-策略蒸馏框架。ATLAS将指南证据构建为用药安全图,通过针对性问题更新患者状态,并将相关关系蒸馏为患者特异性药物冲突图(PMCG)。基于风险优先的多智能体策略利用PMCG筛查禁忌症、评估注意事项与监测需求、识别更安全的替代方案并验证最终用药方案。本文还提出GeriMedBench,这是一种用于测试安全关键信息获取与循证决策修正的交互式基准。在欧洲非交互式多病基准、亚洲交互式多病基准及亚洲非交互式跨指南基准上,ATLAS在对比系统中实现了最强的完整决策性能。在欧洲非交互式多病基准上,其严格成功率(Strict Success Rate)较最强的专有LLM基线高出53.73个百分点,整体安全推理得分(OSRS)高出14.63分,且在自动评估器下无不安全建议。盲法临床医生评估显示,ATLAS在所有五项标准上的平均评分更高,且仅在1个ATLAS案例中发现潜在不安全建议,而Gemini案例中有2个。

英文摘要

Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on conditions, medications, and geriatric risks that users may omit. We introduce ATLAS, a coupled graph--policy distillation framework for patient-adaptive medication safety. ATLAS structures guideline evidence as a medication-safety graph. Targeted questions update the patient state and distill relevant relations into a patient-specific medication conflict graph (PMCG). A risk-first multi-agent policy uses the PMCG to screen contraindications, assess cautions and monitoring needs, identify safer alternatives, and verify the final medication plan. We also introduce GeriMedBench, an interactive benchmark that tests safety-critical information acquisition and evidence-based decision revision. Across a European non-interactive multimorbidity benchmark, an Asian interactive multimorbidity benchmark, and an Asian non-interactive cross-guideline benchmark, ATLAS achieves the strongest complete-decision performance among the compared systems. On the European non-interactive multimorbidity benchmark, it exceeds the strongest proprietary LLM baseline by 53.73 points in Strict Success Rate and 14.63 points in overall safety reasoning score (OSRS), with no unsafe recommendations under the automated evaluator. A blinded clinician evaluation gives ATLAS higher mean ratings across all five criteria and flags potentially unsafe recommendations in one ATLAS case and two Gemini cases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑