arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GuideSkill:面向基于指南的临床推理的可执行大语言模型智能体技能演化框架

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

Lang Cao, Yuhao Shen, Tianyang Luo, Simo Du, Hao Peng, Yue Guo

arXiv 2607.26160首次发表:更新:

AI 中文总结

本文提出与模型无关的GuideSkill框架,通过可执行技能结合指南流程与病例诊断模式,在多基准和骨干模型上显著提升临床推理性能,技能可靠且实用。

AI 中文摘要

临床实践指南(CPGs)包含诊断标准,但大语言模型(LLM)系统通常仅检索指南文本或通过训练吸收相关内容,而非执行其规则。本文提出GuideSkill,这是一个外部推理层,可将特定疾病的标准编译为返回有序诊断支持分数的可执行函数。GuideSkill-Zero从指南初始化,而GuideSkill-Evo利用病例-诊断对优化覆盖的技能并补充缺失的诊断。推理时,LLM会提出鉴别诊断,匹配每个技能所需的特征,并将其排序与执行的技能分数融合。在四个基准和四个骨干模型上,GuideSkill-Zero的宏平均准确率较基于指南的检索增强生成(RAG)平均提升13.45%;GuideSkill-Evo在所有骨干模型上均达到最高宏平均,较直接推理相对提升18.49%,并将金标准技能覆盖率从56.5%提升至99.5%。在Qwen3.5-9B上,它还在不更新骨干模型的情况下,较最强的参数更新基线提升11.16%。专家评估进一步表明,GuideSkill生成的技能具有临床合理性且被广泛接受,说明其初始化和演化的规则可靠且具有实际意义,这些结果支持可执行技能作为一种与模型无关的机制,用于结合指南衍生的流程与病例衍生的诊断模式。

英文摘要

Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scores. GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case--diagnosis pairs to refine covered skills and add missing diagnoses. At inference, an LLM proposes a differential diagnosis, grounds the features required by each matched skill, and fuses its ranking with the executed skill scores. Across four benchmarks and four backbones, GuideSkill-Zero improves macro-average accuracy over guideline RAG by 13.45% on average. GuideSkill-Evo achieves the highest macro-average for every backbone, improves over direct inference by 18.49% relatively, and increases gold-label skill coverage from 56.5% to 99.5%. On Qwen3.5-9B, it also exceeds the strongest parameter-update baseline by 11.16% without updating the backbone. Expert evaluation further indicates that GuideSkill produces clinically sound and broadly acceptable skills, suggesting that its initialized and evolved rules are reliable and practically meaningful. These results support executable skills as a model-agnostic mechanism for combining guideline-derived procedures with case-derived diagnostic patterns.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑