EHR-RobustGym:用于鲁棒临床推理的智能体基准测试与训练
EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning
- Zhejiang University(浙江大学)
- Ant Healthcare, Ant Group(蚂蚁集团蚂蚁医疗健康)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对EHR噪声导致临床智能体推理不鲁棒的问题,提出EHR-RobustGym环境,基于MIMIC-IV构建5,486个干净-噪声对,通过微调和强化学习提升鲁棒性,并泛化至外部基准。
AI中文摘要:
在医院工作流程中,电子健康记录(EHR)通常包含噪声,可能缺乏确认临床查询中所引用事件或测量结果所需的证据。即使数据库检索成功,临床智能体也可能忽略此类差异,返回看似合理但缺乏支持的答案。我们引入了EHR-RobustGym,一个可扩展且交互式的环境,用于评估和训练基于噪声EHR的鲁棒临床智能体。EHR-RobustGym基于MIMIC-IV医院记录(365K名患者、31张表、超过5亿条记录)构建,包含5,486个干净-噪声对,涵盖六种临床意图以及患者级和人群级查询。这些配对测试了对记录级、数值级和查询级噪声的鲁棒性,同时交互式SQL/Python执行和结果验证支持轨迹收集和训练。对多个LLM的评估揭示了显著的鲁棒性差距:专有模型和大规模开放权重模型在干净问题上的平均任务成功率从62.2%下降到噪声问题上的37.9%。在k=4时,大多数评估模型的pass^k一致性低于50%,暴露了临床任务完成中的不稳定性。在EHR-RobustGym中进行监督微调和强化学习可提高性能,其收益可泛化到五个外部EHR基准。这些结果共同将EHR-RobustGym定位为评估和改进临床智能体证据基础鲁棒性的测试平台。
英文摘要:
In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.