发表机构
Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出CLINLENS基准,包含200项基于5种MIMIC资源的临床可执行任务,实验显示现有编码智能体与生物医学系统在正确临床分析上存在显著差距。
AI 中文摘要
临床数据科学智能体必须将异构纵向记录转化为可审计的分析结果,然而现有基准大多孤立地处理医学问答、结构化表格推理或通用科学知识库。我们推出CLINLENS,这是一个基于5个关联MIMIC资源的200项可执行任务的基准,涵盖结构化电子健康记录、病历、心电图、胸部X光片及超声心动图。4×5分类体系将4种患者时间范围与5种分析能力交叉对应。以程序为核心的逆向合成法,将每个受限半原始数据包与评估者私有的参考工作流程配对,检查所需工件、队列和时间语义及最终答案。在固定的126项任务套件中,24种标准化模型支架配置中性能最强的模型,尽管EXECSUCCESS达到100%,但scope-macro STRICTPASS仅达到56.3%。作为参考,单独配置的编码智能体可解决126项任务中的83项,而适配GPT-4o-mini的5个生物医学系统的scope-macro STRICTPASS最多仅为2.9%。这些结果揭示了可执行提交与正确临床分析之间存在巨大差距。
英文摘要
Automating clinical data science requires agents to translate patient-centric objectives into executable workflows grounded in longitudinal multimodal evidence. Existing benchmarks provide limited coverage of two essential capabilities: grounding analyses in the correct patient episode and harmonizing multimodal observations. We introduce ClinLens, a patient-centric benchmark for evaluating large language model (LLM) agents on long-horizon clinical data science workflows. ClinLens comprises 126 executable tasks spanning five data modalities, four analytical scopes, and five analytical capabilities. The tasks assess agents' ability to maintain a coherent patient-specific analytical context across interdependent stages of evidence grounding, multimodal integration, and analytical pipeline generation and execution. We evaluate general-purpose, coding, and biomedical agents, with task runs averaging 36.93 steps and 8.06 minutes. The strongest evaluated agent achieves an execution success rate of 86.4% and a final-answer accuracy of 54.0%, but a strict pass rate of only 29.4%. Failure analysis identifies patient-episode grounding as the most frequently recorded failure category under the evaluation procedure. By jointly evaluating analytical outcomes and the validity of the supporting clinical context, ClinLens provides a testbed for advancing reliable agents for longitudinal clinical data science.