arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedicalHarness:医疗任务上LLM与智能体框架的受控评估

MedicalHarness: A Controlled Evaluation of LLMs and Agent Harnesses on Medical Tasks

Ziqing Wang, Lili Zhao, Kaize Ding

arXiv 2610.05778首次发表:更新:

发表机构

Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MedicalHarness,通过受控基准和框架分解,系统评估LLM智能体框架对医疗任务性能的影响,发现框架与模型交互约占结果方差四分之一,且无单一框架普遍最优。

AI 中文摘要

LLM智能体越来越多地被构建用于医疗工作,并在临床基准上进行评分。然而,每个这样的分数都来自运行在智能体框架内的模型,该框架是控制模型与其环境之间循环的系统。因此,智能体的分数是模型-框架对的属性。对于医疗智能体,结果随框架变化多少很少被测量。测量这种变化并解释它,带来了两个挑战。首先,框架比较必须只改变框架,并在多个模型和任务类型上重复进行。其次,比较整个框架会将它们的机制捆绑在一起,因此无法显示单个机制何时有帮助。为了解决这些挑战,我们提出了MedicalHarness,一项关于医疗任务上模型和智能体框架的受控研究。我们首先构建了MedicalHarnessBench,以评估智能体在四个领域的107个任务上的表现,每个领域测试不同的框架能力。使用该基准,我们在五个智能体框架下运行五个开放权重模型,在比较中仅改变框架,并分析结果和执行轨迹。为了研究单个机制,我们构建了MH-Lab,一个受控框架,在共享执行循环中一次关闭上下文管理、规划或工具暴露之一。我们发现,框架及其与模型的交互约占结果方差的四分之一,并且没有单一框架在所有模型和任务上都是最佳的。代码和数据可在该https URL获取。

英文摘要

LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is therefore a property of a model--harness pair. For medical agents, how much outcomes change with the harness has rarely been measured. Measuring this change, and explaining it, raises two challenges. First, a harness comparison must change nothing but the harness and be repeated across models and kinds of task. Second, comparing whole harnesses leaves their mechanisms bundled together, so it cannot show when an individual mechanism helps. To address these challenges, we present MedicalHarness, a controlled study of models and agent harnesses on medical tasks. We first build MedicalHarnessBench to evaluate agents on $107$ tasks across four domains that each test a different harness capability. Using this benchmark, we run five open-weight models under five agent harnesses, changing only the harness within a comparison, and analyze both outcomes and execution traces. To study individual mechanisms, we build MH-Lab, a controlled harness that switches off context management, planning or tool exposure one at a time within a shared execution loop. We find that the harness and its interaction with the model account for about a quarter of the outcome variance, and that no single harness is best across models and tasks. Code and data are available at https://github.com/REAL-Lab-NU/MedicalHarness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑