Asclepius:面向长时程临床智能体的自适应调控框架
Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
浏览论文内容
中文总结 AI 辅助
针对长时程临床智能体执行失败问题,提出Asclepius自适应调控框架,通过自我进化手册、技能库和子智能体分区,将关键行动正确性提升22%-25%,并保持诊断准确性。
中文摘要 AI 辅助
大语言模型智能体主要在短时程、单任务轨迹上进行基准测试,然而实际部署中,智能体需在资源竞争条件下运行数小时,这暴露出另一类不同的失败模式。我们使用临床环境模拟器(CES)作为测试平台,在该模拟器中,智能体需在持续的时间和资源压力下管理整个急诊科轮班:在结构化、多维度的评分体系下,长时程执行失败可在单次运行中可量化地显现。在CES上,现有智能体在大多数情况下能得出正确诊断,但未能提供完整且及时的关键行动,揭示了执行差距。我们将这一差距归因于三种长时程失败模式,每种模式均被操作化为逐轨迹计数器:指令遵循漂移、治疗不完整以及严重程度公平性方面的及时性差距。随后,我们引入Asclepius,一种自适应智能体脚手架,具有自我进化的调控框架,该框架在轮班之间根据轨迹级反馈重写操作手册,并配备外部化临床技能库以存储高风险治疗方案知识,以及三个隔离的子智能体,用于在患者队列中划分每轮决策。在调控框架进化过程中从未观察到的保留批次上,Asclepius将关键行动正确性提高了22%(p = 0.024),相较于强基线智能体框架,同时保持了诊断准确性,在来自三个模型族的五个大语言模型评判者上均取得一致提升;在完整的十批次集合上,关键行动的改进达到25%,及时性改进达到13%。这三种失败模式构成耦合瓶颈:只有当所有三个组件协同作用时,才会出现决定性的改进。
英文摘要
LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional grading. On CES, current agents reach the correct diagnosis in most cases yet fail to deliver complete and timely critical actions, revealing an execution gap. We attribute this gap to three long-horizon failure modes, each operationalized as a per-trace counter: instruction-adherence drift, treatment incompleteness, and a severity-equity gap in timeliness. We then introduce Asclepius, an adaptive agent scaffolding with a self-evolving harness that rewrites the operating manual between shifts from trace-level feedback, an externalized clinical skills library for high-stakes regimen knowledge, and three isolated subagents that partition per-turn decisions across the patient queue. On held-out batches never observed during harness evolution, Asclepius improves critical-action correctness by 22% (p = 0.024) over a strong baseline agent framework while preserving diagnostic accuracy, with consistent gains across five LLM judges from three model families; on the full ten-batch set, improvements reach 25% on critical actions and 13% on timeliness. The three failure modes form a coupled bottleneck: decisive reductions appear only when all three components act together.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- Harvard Medical School(哈佛医学院)
机构由 AI 辅助整理,请以论文原文为准。