发表机构
Mass General Brigham; University of Edinburgh; University of Oxford; Stanford University; Harvard Medical School(麻省总医院布莱根医疗体系; 爱丁堡大学; 牛津大学; 斯坦福大学; 哈佛医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文综述了医疗领域大型语言模型评估的四个关键领域:研究设计、统计方法、能力评估和临床背景评估,强调评估方法需与研究问题对齐,为严格评估提供实用基础。
AI 中文摘要
大型语言模型(LLMs)在医学领域的应用日益广泛,对其进行评估对于确保其带来益处而非伤害至关重要。由于多种原因,这种评估可能比传统机器学习更具挑战性,包括概率性和开放式输出,以及行为随提示设计和累积上下文而变化。本综述涵盖LLM评估的四个关键领域:研究设计原则、统计方法、能力评估和临床背景评估。能力评估考虑不同的基准,包括多项选择、智能体(agentic)和多轮基准,以及诸如令牌使用等操作指标。临床背景评估涉及确定自由文本输出的准确性,如人工审查和LLM作为评审(LLM-as-a-judge),以及临床试验方法。在各章节中,我们描述了基本概念和潜在陷阱,同时强调将评估方法与研究问题对齐的重要性。总之,本文旨在为设计和执行对医疗保健LLM的严格评估提供实用基础。
英文摘要
Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.