arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

未知并非正常:将语言模型提取与基于规则的决策逻辑分离用于临床风险评分

Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores

Nicolás Vera Zúñiga

arXiv 2609.34112首次发表:更新:

AI 中文总结

本研究提出将LLM三态提取与确定性决策逻辑分离,仅询问可改变决策的问题,在临床风险评分中达到同等准确率且提问减半,避免低估分诊,并支持小型本地模型。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用于从自由文本笔记中计算临床风险评分。笔记往往不完整,将未记录的结果视为正常可能会静默地错误分类患者。我们测试了将三态提取(存在、不存在或未知,由LLM执行)与决策逻辑(确定性代码,在未知输入上计算分数界限)分离,是否能让系统只提出那些可能改变决策的问题。在跨越六个计算器(HEART、CURB-65、qSOFA、PERC、Wells、Cockcroft-Gault)的1,200个合成急诊病例中,通过模拟临床医生回答问题,我们将这种界限策略与询问每个缺失输入、缺失即正常模式以及端到端LLM代理(Claude Opus 5.5)进行了比较。使用Claude Haiku 4.5作为提取器时,界限策略在准确率上与询问所有输入相匹配(99.4%对99.4%),但问题数量减半(每例0.92对1.78个),且没有无关问题。将缺失视为正常导致准确率降至91.2%,并导致8.5%的患者被低估分诊(95%置信区间7.1-10.2),这种低估分诊在杂乱笔记和嘈杂临床医生的情况下持续存在。代理在理想条件下同样准确(99.6%),但其9.5%的问题无关;在嘈杂临床医生情况下,其准确率低于界限策略(83.5%对87.0%,p<0.001),并在2.7%的病例中过早提交(界限策略为0%)。一个9B参数的本地模型作为提取器达到了神谕级准确率(99.8%)。在来自MedCalc-Bench的584份真实病例报告中,只有52%包含足够信息来确定类别(HEART为13%)。通过明确推理未知数的代码来路由决策,避免了过早提交和无关问题,将提问数量减半,并且适用于小型本地模型。

英文摘要

Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision logic (deterministic code computing score bounds over unknown inputs) lets a system ask only questions that can change the decision. On 1,200 synthetic emergency cases across six calculators (HEART, CURB-65, qSOFA, PERC, Wells, Cockcroft-Gault), with a simulated clinician answering questions, we compared this bounds policy with asking for every missing input, a missing-equals-normal schema, and an end-to-end LLM agent (Claude Opus 5.5). With Claude Haiku 4.5 as extractor, the bounds policy matched ask-all accuracy (99.4% vs 99.4%) with half the questions (0.92 vs 1.78 per case) and no irrelevant ones. Treating missing as normal dropped accuracy to 91.2% and under-triaged 8.5% of patients (95% CI 7.1-10.2), and under-triage persisted under messy notes and a noisy clinician. The agent was equally accurate under ideal conditions (99.6%) but 9.5% of its questions were irrelevant; with a noisy clinician it was less accurate than the bounds policy (83.5% vs 87.0%, p<0.001) and committed prematurely in 2.7% of cases (bounds: 0%). A 9B local model as extractor reached oracle-level accuracy (99.8%). In 584 real case reports from MedCalc-Bench, only 52% contained enough information to determine the category (HEART 13%). Routing decisions through code that reasons explicitly about unknowns avoids premature commitment and irrelevant questions, halves the questions asked, and works with small local models.

Comments14 pages (7 of main text), 5 figures, 2 tables, appendix included; full supplementary material in the code repository. Code: https://github.com/nicoveraz/calc-bounds (archived: https://doi.org/10.5281/zenodo.23004726)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑