arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPAgentBench-2K:复杂临床动作空间中的大语言模型智能体基准测试

GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space

Boqi Chen, Xudong Liu, Yunke Ao, Heejin Do, Jianing Qiu

arXiv 2608.30188首次发表:更新:

发表机构

ETH Zurich; MBZUAI; University of Toronto(苏黎世联邦理工学院; 穆罕默德·本·扎耶德人工智能大学; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出首个面向初级保健临床决策的CMDP LLM智能体基准GPAgentBench-2K,评估16种LLM时发现其存在临床质量-安全差距,且C-GRPO建模约束仍未达临床安全要求。

AI 中文摘要

大语言模型(LLMs)作为临床智能体展现出巨大潜力,但现有基准测试将临床工作流程简化为静态预测或具有粗糙动作集的无约束马尔可夫决策过程(MDP)。为解决这一问题,我们推出GPAgentBench-2K——首个面向初级保健临床决策的约束马尔可夫决策过程(CMDP)LLM智能体基准,其构建基于经专家验证的真实全科医生(GP)接诊记录。该环境涵盖六种基础临床动作的全范围,对动作空间施加拓扑工作流程先验,并将安全导向的弃权(不执行)作为一级结果。对16种最先进LLM的评估显示,随着动作空间规模扩大,性能显著下降。关键是,我们发现了临床质量-安全差距:即便诊断准确率最高的前沿模型,在超过一半的高风险案例中也会违反安全约束。最后,我们使用约束分组相对策略优化(C-GRPO)建立了参考点,结果显示,虽然显式建模约束比无约束强化学习方法的性能有所提升,但仍远未达到临床可接受的安全水平。

英文摘要

Large Language Models (LLMs) show great potential as clinical agents, yet existing benchmarks reduce clinical workflows to static predictions or unconstrained Markov Decision Processes (MDPs) with coarse action sets. To address this, we introduce GPAgentBench-2K, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters. Our environment models a full spectrum of six foundational clinical actions, imposes a topological workflow prior over the action space, and operationalizes safety-informed abstention as a first-class outcome. Evaluating 16 state-of-the-art LLMs reveals a significant performance degradation as the action space scales. Crucially, we uncover a clinical quality-safety gap: even frontier models with the highest diagnosis accuracy violate safety constraints in over half of high-risk cases. Finally, we establish a reference point using Constrained Group Relative Policy Optimization (C-GRPO), and show that while explicitly modeling constraints improves performance over unconstrained RL methods, it remains far from clinically acceptable safety.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑