arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向医疗NLP智能体的工作流感知基准测试

Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

Junyi Yao, Baichuan Li, Zihao Zheng, Jiayu Long

arXiv 2609.00296首次发表:更新:

发表机构

Washington University in St. Louis; Southern Methodist University(圣路易斯华盛顿大学; 南卫理公会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对医疗NLP智能体现有评估局限,提出回合级工作流感知评估协议,设四项任务模板,填补静态基准与工作流研究间的评估空白。

AI 中文摘要

大型语言模型(LLM)智能体越来越多地被用于医疗任务,如临床文档生成、证据检索、患者消息发送及护理协调。但许多评估仍局限于静态医学问答或单次生成,未能充分体现纵向状态、中断及人工交接的情况。我们提出一种针对医疗NLP智能体的回合级评估协议,该协议将证据划分为模型、智能体及模拟工作流行为三类,规定了五字段回合模式,并定义了状态连续性、证据可追溯性及升级决策的标注与评分规则。该协议实例化为文档更新、证据检索、患者消息发送及分诊交接四项任务模板,不声称衡量临床结果或部署价值,而是在静态基准与前瞻性工作流研究之间提供可复现的中间评估层,对遗漏与不必要的升级采取显式成本敏感处理。

英文摘要

Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical question answering or one-shot generation, under-representing longitudinal state, interruptions, and human handoffs. We introduce an episode-level evaluation protocol for healthcare NLP agents. The protocol separates evidence across model, agent, and simulated-workflow behavior; specifies a five-field episode schema; and defines annotation and scoring for state continuity, evidence traceability, and escalation decisions. It is instantiated as four task templates: documentation update, evidence retrieval, patient messaging, and triage handoff. The protocol does not claim to measure clinical outcomes or deployment value. Instead, it supplies a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.

Comments4 pages, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑