临床大语言模型中证据充分性提示的依赖判断的安全增益和特定模型的有用性成本
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
浏览论文内容
中文总结 AI 辅助
研究临床大语言模型中证据充分性提示的安全增益及有用性成本,通过在公共数据基准中让四个模型用标准提示和包装器回答问题,发现安全增益有方向和幅度差异且依赖评判者,同时存在模型特定的有用性成本。
中文摘要 AI 辅助
背景:大语言模型(LLM)评判者越来越多地对临床语言模型在证据不完整时给出过度自信答案的情况进行评分,然而,所衡量的“安全增益”是反映实际行为变化还是评判者的校准仍未解决。以结构化证据充分性提示作为测试案例,我们探讨它是否能减少不安全的过度自信答案,这种效果在多大程度上依赖于评分评判者,以及它在有用性方面的成本。方法:在回顾性公共数据基准(Real-POCQi、HealthBench、MedRBench)中,四个模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash、Grok 4.3)用标准提示和包装器回答了一个完全配对的通用面板(1200个单元格)。预先指定的终点是主要评判者(GPT-5.4-nano)评分的不安全过度自信的配对减少;二次分析增加了不同家族的评判者(Claude Sonnet 5)、正确性评判者、匹配的支架对照和盲法三临床医生评审。结果:不安全的过度自信从49.3%降至24.7%,配对减少24.7个百分点(95%置信区间21.8-27.7;p<0.001),在模型和释义中方向稳健。幅度依赖于评判者:Sonnet在方向上一致,但效果几乎减半(+13.1个百分点),存在单向分歧。盲法临床医生将主要评判者描述为高敏感性(1.00)、低特异性(0.55)的筛查,而非校准率。增益带来了特定模型的有用性成本(正确诊断从80.3%降至50.3%):GPT-5.5几乎无成本,Gemini几乎完全有成本(-58分)。匹配的支架对照显示了真正的行为变化,而非评判者循环。结论:LLM评判的临床安全效果应以方向和相对的方式报告,以人类评审为锚定,并与有用性联合评估,而不是作为校准的绝对率。这并未确定临床部署的准备情况。
英文摘要
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.