arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MediSkill-Evo:面向证据依据的临床交互的过程约束自进化

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao

arXiv 2608.23397首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; Beijing Institute of AI Safety and Governance; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所; 北京人工智能安全与治理研究院; 中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MediSkill-Evo是无需主干微调的临床智能体,通过四类经验存储库与过程约束偏好工具优化,在多维度评估中提升了诊断准确率等指标,验证了其临床交互性能。

AI 中文摘要

交互式临床智能体必须在部分可观测的情况下收集决定性证据,并将其转化为有依据的行动。仅最终诊断正确并不能表明智能体遵守了证据和护理流程约束。我们提出MediSkill-Evo,一种无需主干微调即可进化受控流程知识的临床智能体,它将经验分为临床技能、流程规则、符号模式和测量程序四类存储库。出处、支持度、重放以及控制器定义的安全检查共同约束内容发布至冻结的测试时快照。过程约束偏好工具将证据与其来源绑定,拒绝控制器无效的候选,并以安全优先的临床流程评判器对行动进行排名。我们在相同的医生回合限制下,针对两个主干端点和六个受控压力维度评估完整智能体系统。在300个保留的Qwen交互案例中,与AgentClinic相比,MediSkill-Evo将诊断准确率从61.33%提升至69.00%,治疗意图覆盖率从33.62%提升至66.44%,同时将自动评分的严重故障从31.00%降至16.33%。在源自30个案例的180个硬隔离条件下,患者行为压力下的目标恢复率达到93.61%,时间证据下为100.00%,分诊红旗下为92.22%。一项探索性的100例MedSAM对比评估了请求门控工具接口的可行性。这些结果为固定评估套件上的完整系统提供了描述性端到端证据,而非针对单个存储库的因果证据,也非自动评判器的临床验证。

英文摘要

Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existing agents struggle to convert experience into reusable process knowledge with explicit provenance and authority. To address this gap, we introduce MediSkill-Evo, which self-evolves governed process knowledge without fine-tuning the backbone. It realizes this self-evolution by updating clinical, process, symbolic, and visual knowledge in four typed banks under type-specific validation and scope rules. The Process-Constrained Preference Harness then turns validated knowledge into action by grounding candidates in evidence and prioritizing safer decisions. We evaluate on 300 MIMIC-IV-derived FullChain encounters, 180 hard-isolation conditions covering six process obligations, and 100 multimodal NEJM image-diagnosis cases. On Qwen FullChain, MediSkill-Evo improves diagnosis accuracy by 7.81% and treatment-intent coverage by 70.67% over the best-performing prior agent, while reducing critical failures by 43.04%. Under stress, it improves the stress-process composite by 7.77% and required-action completion by 12.41% over the best-performing agent for each metric, with stronger patient-fact, temporal-evidence, and triage-red-flag recovery and no controller-scored errors in unavailable-evidence, treatment, and triage safety checks. On multimodal NEJM diagnosis, MediSkill-Evo with optional MedSAM localization improves diagnosis accuracy by 2.56% and core score by 18.96% over the best-performing memory agent. Code is available at https://anonymous.4open.science/r/mediskill-evo_anonymous-68E7.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑