按构造可验证:临床问答中逐字引用的声明级评估
Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
浏览论文内容
中文总结 AI 辅助
本文评估了十二个大型语言模型在临床问答中端到端生成逐字引用并完全证实声明的能力,发现多数模型虽能附加引用,但完全证实率低,揭示了可验证临床QA系统的能力差距。
中文摘要 AI 辅助
大型语言模型(LLMs)已被广泛用于临床问答(QA)。当前系统可以为答案附加引用,但这些引用往往指向宽泛的文本,使时间紧迫的临床医生无法高效验证。另一种方案是确保回答在构造上即可验证:提供来自参考材料的细粒度逐字引用,以证实各项声明,使用户无需打开其他文档即可验证答案。在本文中,我们评估了当前模型端到端执行此任务的能力:从为每个事实性声明提供引用,到生成逐字引用,再到确保这些引用完全证实声明。为此,我们基于四份临床实践指南构建了一个标准化测试框架,并在222个合成临床问题上评估了十二个LLM,分别测量每个阶段的表现。我们发现,大多数模型仅通过提示即可为超过90%的声明附加逐字引用,但一些轻量级模型(如claude-haiku-4.5)除外。然而,这些引用往往未能完全证实其所伴随声明的每个细节。例如,claude-opus-5为其98.0%的声明生成了逐字引用,但仅完全证实了37.1%。我们的工作揭示了当前LLM在构建可验证临床QA系统方面的能力差距,并为未来研究提供了相关工具。
英文摘要
Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems can attach citations to their answers, but these often point to broad texts, leaving time-pressed clinicians unable to verify them efficiently. An alternative is to ensure that responses are verifiable by construction: providing fine-grained verbatim quotes from reference material that substantiate claims, so users can verify an answer without opening other documents. In this paper, we evaluate the ability of current models to perform this task end-to-end: from providing citations for every factual claim, to producing verbatim quotes, to ensuring that those quotes fully substantiate the claims. To do so, we build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring each of these stages separately. We find that most models can attach verbatim quotes to over 90% of their claims from prompting alone, apart from some lightweight models such as claude-haiku-4.5. Yet these quotes often fail to substantiate every detail of the claims they accompany. For instance, claude-opus-5 produces verbatim quotes for 98.0% of its claims, but fully substantiates only 37.1%. Our work provides insights into the current capability gap of LLMs in building verifiable clinical QA systems, along with artifacts for future research.
发表机构
- Johns Hopkins University(约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。