arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31908cs.AIcs.CL

利用嵌入式编码改进大语言模型的医学计算

Improving Medical Calculation of LLMs with Embedded Coding

Tianshi Ming, Yingying Zhang, Xian Wu

AI总结:

本文提出MedCode框架,通过训练LLMs生成嵌入式可执行代码,将算术操作委托给确定性解释器,在医学计算任务上取得20-30个百分点的绝对准确率提升。

AI中文摘要:

大语言模型(LLMs)在医学考试和问答基准测试中表现良好,但在需要精确数值输出的医学计算任务上仍不可靠。这些计算支持高风险的决策,如药物剂量、器官功能评估和预后评分,即使是小错误也可能导致严重的临床后果。我们提出了MedCode,一个通过训练LLMs生成嵌入式可执行代码来改进医学计算的框架。给定临床上下文,模型识别相关的计算器,提取其输入变量,并生成一个将算术操作委托给确定性解释器的脚本。执行脚本返回计算值以及解释和适当的单位。我们从MedCalc基准构建了监督微调(SFT)和偏好数据集,并额外为重症监护病房(ICU)场景中的计算任务整理了一个数据集。我们进一步提出了加权直接偏好优化(wDPO),它自适应地强调模型难以区分的偏好对。使用LLaMA3-8B、Qwen2.5-7B和Mistral-7B进行的实验显示,绝对准确率提高了20至30个百分点,证明了嵌入式代码生成在医学计算中的有效性。

英文摘要:

Large Language Models (LLMs) perform well on medical examinations and question-answering benchmarks, but remain unreliable on medical calculation tasks that require exact numerical outputs. These calculations support high-stakes decisions such as medication dosing, organ-function assessment, and prognostic scoring, for which even small errors can have serious clinical consequences. We introduce MedCode, a framework that improves medical calculation by training LLMs to generate embedded executable code. Given a clinical context, the model identifies the relevant calculator, extracts its input variables, and produces a script that delegates arithmetic operations to a deterministic interpreter. Executing the script returns the calculated value together with an explanation and the appropriate unit. We construct supervised fine-tuning (SFT) and preference datasets from the MedCalc benchmark and additionally curate a dataset for calculation tasks in Intensive Care Unit (ICU) scenarios. We further propose weighted Direct Preference Optimization (wDPO), which adaptively emphasizes preference pairs that are difficult for the model to distinguish. Experiments with LLaMA3-8B, Qwen2.5-7B, and Mistral-7B show absolute accuracy gains of 20--30 percentage points, demonstrating the effectiveness of embedded code generation for medical calculation.

↑