arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25485cs.AIcs.CL

PatientAgentBench:一个用于评估面向患者的健康人工智能代理的基准框架

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

  • Amazon Health AI(亚马逊健康人工智能)

机构由 AI 辅助整理,请以论文原文为准。

Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, … 展开作者

Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf

AI总结:

研究面向患者的健康人工智能代理评估问题,核心方法是用PatientAgentBench框架,通过语言模型评审团在多维度评估,主要贡献是发现模型临床差距,发布可重复、经临床医生验证的评估标准助领域弥补差距。

AI中文摘要:

健康人工智能正在从回答问题向与患者对话、推理健康记录并代表患者行动的智能系统发展。初级保健可防范诊断错误和不安全护理;该领域的辅助代理需要针对相同风险进行评估。当前基准侧重于通过孤立的问答或面向临床医生的任务评估医学知识。PatientAgentBench对面向患者的智能医疗保健进行基准测试;它评估一个基础模型,该模型封装在一个带有医疗保健工具沙盒的代理中,与模拟患者进行对话。每次对话由一个作为评审团的语言模型通过一百多个与对话无关、基于临床医生的标准在六个维度上进行评分。为了验证一致性,持牌临床医生对共享对话进行注释,评审团和专家评分者之间的相邻一致性为79-93%,与临床医生评分者之间的一致性相当或更高。我们在相同的1200个场景中对四个系列的10个模型进行了基准测试,发现了临床差距。分诊质量是最具区分性的维度:最弱模型的通过率从32%上升到最强模型的88%,代理经常根据行政请求行事而不进行临床筛查。临床安全性和工作流程准确性遵循相同模式:最弱的模型经常失败,编造未执行的行动,而前沿模型仅在1-3%的情况下失败,原因是工具输出未经验证以及紧急情况下遗漏危机资源。能力更强的模型缩小了这些差距,但并未消除;最强的模型在总体5分中仅得4.25分。这些失败仅在针对真实患者记录的持续、使用工具的对话中出现,证实随着医疗保健智能系统获得自主性,静态基准是不够的。我们将该框架作为可重复、经临床医生验证的评估标准发布,以帮助该领域弥补这一差距。

英文摘要:

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.

补充信息

相关深度报道

↑