发表机构
MIT Critical Data; University of Pavia; Humanitas University(麻省理工学院关键数据; 帕维亚大学; 人文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大模型算术不可靠问题,提出Program-Solve接口:模型编写Python代码由本地执行器确定性求解,在MedCalc-Bench上32B模型提升显著(+7.05%),但7B提升不显著,且不能替代验证公式。
AI 中文摘要
大型语言模型在算术方面不可靠,这在临床计算器中是一个问题,因为单个数值错误就会改变推荐结果。标准的应对方法是逐个将每个计算器硬编码为经过验证的函数。我们测试了一种替代方案:模型不进行计算,而是编写针对具体病例的Python代码,由受限的本地执行器作为确定性求解器运行,模型的任务简化为决定如何使用它。我们在MedCalc-Bench Verified(1100个病例,55个计算器)上评估了这种Program-Solve接口,与直接模型算术和手写的22个计算器库进行比较,使用Qwen2.5-7B和Qwen2.5-32B-AWQ,并对照当前临床指南审计了基准的公式,标记了55个中的16个存在版本、使用或系数问题。在提供公式和黄金变量且两种途径都读取整个病历的情况下,对于7B模型,将计算交给求解器并非可靠的改进(75.31%对比72.02%,配对+3.29个百分点,95%计算器簇区间为[-3.49, 10.38]),但对于32B模型则是(90.53%对比83.47%,+7.05 [0.47, 14.60],明确大于零)。手写库在其支持的440个病例上是精确的,但在其他情况下弃权(不执行)(总体40.0%)。因此,在匹配的公式、变量和病历访问条件下,添加执行器对某些开放权重模型的帮助大于其他模型,并且无论如何都不能替代经过验证的公式或可靠的变量提取。
英文摘要
Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the model does not calculate. Instead, it writes case-specific Python that a restricted local executor runs as a deterministic solver, and the model's task reduces to deciding how to use it. We evaluate this Program-Solve interface on MedCalc-Bench Verified (1,100 cases, 55 calculators) against direct model arithmetic and a hand-written 22-calculator library, using Qwen2.5-7B and Qwen2.5-32B-AWQ, after auditing the benchmark's formulas against current clinical guidelines and flagging 16 of 55 with version, use or coefficient concerns. With formulas and gold variables supplied and both routes reading the whole note, handing off to the solver is not a reliable advantage at 7B (75.31% against 72.02%, a paired +3.29 points with a 95% calculator-cluster interval of [-3.49, 10.38]) but is one at 32B (90.53% against 83.47%, +7.05 [0.47, 14.60], clear of zero). The hand-written library is exact on its 440 supported cases but abstains elsewhere (40.0% overall). Adding an executor thus helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.