同一患者,不同医嘱:临床LLM智能体在重复运行下的动作级可靠性
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
- Florida International University(佛罗里达国际大学)
- Boston University(波士顿大学)
- Stevens Institute of Technology(史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究通过“相同输入重跑”方法揭示临床LLM智能体在重复运行中产生动作级分歧且未被基准分数记录,提出重复运行评估与动作级稳定性报告的必要性。
中文摘要 AI 辅助
一个临床智能体基准可以在相同输入上报告相同的结论,而智能体在每次运行中却提交了实质不同的医嘱。这类智能体会开具检查、请求药物和转诊,但基准通常每次任务只评分一次运行,很少询问相同输入是否产生相同动作;我们使用的基准MedAgentBench只对单次尝试评分并如此说明。为衡量这一差距,我们引入了“相同输入重跑”,即固定所有输入重放任务,并比较医嘱而非分数,采用六项可靠性指标,并将其应用于来自五个可写功能族的50个任务的1000次MedAgentBench运行,使用两个低于百亿参数、量化至四比特的开源模型,以及两个温度设置。该研究确立了动作级分歧的存在且可能未被分数记录,而非任何比率具有普遍性。在8B模型、温度0.7下,所有43个医嘱组在五次相同运行中均产生了不同的医嘱集合,26个组在某些运行中发出医嘱而在其他运行中未发出,28个组记录了不同的编码值、剂量或分析物。在这43组中的22组中,基准对实质不同的行为报告了相同的失败结论,正如它对4B模型在0.7温度下的全部10个分歧组所做的那样。医嘱在不同运行中也到达不同的端点,其中一个端点被记录服务器拒绝,而智能体却被告知成功。这些发现推动了临床智能体基准中的重复运行评估、动作级稳定性报告和执行忠实的环境反馈。
英文摘要
A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whether identical inputs produce identical actions; MedAgentBench, the benchmark we use, scores a single attempt and says so. To measure this gap we introduce "same-input rerun", which replays a task with every input held fixed and compares the orders rather than the score, with six reliability metrics, and apply it to 1000 MedAgentBench runs across 50 tasks from its five write-capable families, two open-weight models below ten billion parameters quantised to four bits, and two temperatures. The study establishes that action-level divergence exists and can pass unrecorded by the score, not that any rate generalises. Under the 8B model at temperature 0.7, all 43 ordering groups emit a different set of orders across five identical runs, 26 emit the order on some runs and not others, and 28 record a different coded value, dose or analyte. In 22 of those 43 the benchmark reports the same failing verdict for materially different behaviour, as it does for all 10 divergent groups of the 4B model at 0.7. Orders also reach different endpoints across runs, one of which the record server rejects while the agent is told it succeeded. These findings motivate repeated-run evaluation, action-level stability reporting and execution-faithful environment feedback in clinical-agent benchmarks.